chiprook
← AI
AISeptember 23, 2026, 10:27

Qwen3.8-27B on one RTX 3090 vs two: +20% decode, +14% cold prefill

Tests show Qwen3.8-27B at W4A16 fits on a single RTX 3090 and decodes at 125–155 tokens/sec. Splitting it across two cards with tensor parallelism adds about 20% to decode and 14% to cold prefill, while cached prompts answer 2–3x faster.

Qwen3.8-27B on one RTX 3090 vs two: +20% decode, +14% cold prefill
#Qwen#Nvidia#RTX3090#VLLM
Read next
AI

Experts debate risk of an AI takeover of the internet

AI

Breeze TTS 2 tops ElevenLabs on TTS leaderboard but bans commercial use

AI

DeepSeek V4.1-Flash matches Claude Opus 5 on coding at a fraction of the cost

AI

Jev is a narrow JSON classifier, not a frontier model