chiprook
← AI
AIOctober 3, 2026, 01:07

Qwen3.8-27B runs on a single RTX 3090 via vLLM

syv-ai published a serving setup for Qwen3.8-27B on a single 24 GB consumer GPU with vLLM and an OpenAI-compatible API. Batch mode hits ~1,035 tok/s decode at 64 concurrent requests, while single-user mode reaches 121 tok/s with MTP speculation.

Qwen3.8-27B runs on a single RTX 3090 via vLLM
#Qwen#VLLM#Nvidia
Read next
AI

Qwen3.8-27B on one RTX 3090 vs two: +20% decode, +14% cold prefill

Hardware

NVIDIA Smooth Motion Runs on RTX 30 Series via Mods

Hardware

WiCi One Turns a Single RTX 5090 Into a Wireless GPU Shared Over Wi-Fi 7

AI

Swift 1.5 Qwen3.8-27B: GGUF variant built to stop overthinking