chiprook
← AI
AIOctober 3, 2026, 18:46

133 GB MoE model runs on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

An enthusiast ran Qwen3.8-Flash-Next in NVFP4 (133 GB) on an RTX 5060 with 8 GB by streaming experts from NVMe, hitting 9–12 tokens/s in Open WebUI. The PyTorch + Triton engine was written with Claude and Codex, and output matched the transformers reference.

133 GB MoE model runs on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
#Nvidia#Qwen#OpenWebUI
Read next
AI

ISOM-R2 streams 1,055,402 tokens at 3.24 GB peak VRAM

AI

OrcaSAQ-2 Shrinks 27B Qwen 3.8 AI Into a 12.3GB Local Model

AI

Bonsai 2 27B: ternary weights shrink a 27B model to 5.9 GB

AI

Bonsai 2 27B compresses Qwen3.8 to 5.9GB while keeping 98.2% of quality