133 GB MoE model runs on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
An enthusiast ran Qwen3.8-Flash-Next in NVFP4 (133 GB) on an RTX 5060 with 8 GB by streaming experts from NVMe, hitting 9–12 tokens/s in Open WebUI. The PyTorch + Triton engine was written with Claude and Codex, and output matched the transformers reference.
- Model: 48 layers, 512 experts, top-10 routing, ~1.27 GB of experts per token
- Resident FP8 part uses 6.13 GB VRAM; 63 GB of experts stay on disk
- Speed rose from 2.68 to 11.65 tokens/s after caching and prefetch
- A 2451-token prompt dropped from 71.6 s to 23.3 s
Read next
AI