chiprook
← AI
AIOctober 4, 2026, 18:00

Strata runs a 125B LLM on a gaming PC with 12 GB VRAM

The Strata project by developer Niko1221 lets a 125-billion-parameter Qwen3.8 mixture-of-experts model run on ordinary gaming hardware. VRAM acts as a cache for experts while the full model stays in system RAM and a 29 GB lookup table lives on the SSD. An RTX 5070 with 12 GB of VRAM reaches roughly 50 to 90 tokens per second.

Strata runs a 125B LLM on a gaming PC with 12 GB VRAM
#Qwen#Nvidia#RTX5070
Read next
AI

Qwen3.8-Flash-Next vs Qwen3.8-27B: 125B Parameters, 6B Active — What the Qwen4 Preview Changes

AI

ISOM-R2 streams 1,055,402 tokens at 3.24 GB peak VRAM

Hardware

Gigabyte AI TOP 100 B850 Packs Dual Radeon AI PRO R9700 With 64 GB VRAM

AI

133 GB MoE model runs on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe