Qwen3.8-27B runs on a single RTX 3090 via vLLM
syv-ai published a serving setup for Qwen3.8-27B on a single 24 GB consumer GPU with vLLM and an OpenAI-compatible API. Batch mode hits ~1,035 tok/s decode at 64 concurrent requests, while single-user mode reaches 121 tok/s with MTP speculation.
- Batch mode: ~1,035 tok/s at 64 requests, 1,222 tok/s with int8 layers
- Single-user: 121 tok/s with MTP, 127 tok/s with SPEC=dflash2
- DFLASH_TOKENS=15 gives 382 tok/s when quoting docs but cuts slots from 8 to 4
- First start pulls a 9.5 GB image and ~20 GB model, serves on port 18020
Read next
AI