Strata runs a 125B-parameter model on a 12 GB consumer GPU
The open-source C++ project Strata applies 2-bit quantization (Q2_0/IQ2_XS) to Qwen3.8-Flash-Next, running a 125-billion-parameter model on a single 12 GB consumer GPU. On an RTX 5070 with Ryzen 5 7600 it reaches 94 tok/s generation and 2,650 tok/s prompt read.
- Q2_0 quantization hits 94 tok/s on a 12 GB RTX 5070, IQ2_XS reaches 79 tok/s
- Prompt read at 32K context reaches 2,650 tok/s
- Local API is OpenAI/Anthropic-compatible, only base_url needs changing
- Windows and Linux installers, NVIDIA and AMD supported
Read next
AI