VIDRAFT ships 180B MoE model that runs on a gaming laptop
VIDRAFT released POCKET-Darwin-180B-GGUF, a 4-bit GGUF build of its 180B MoE model Darwin-180B-RSI. Sparse activation (10 of 512 experts, ~3B active parameters) plus llama.cpp SSD streaming lets the 111 GB model run on a laptop with 8 GB VRAM and 32 GB RAM at 4.17 tokens/sec.
- 4-bit quantization shrinks the model from 360 GB to ~111 GB across four GGUF files
- Only 10 of 512 experts activate, about 3B parameters per token
- Runs at 4.17 tokens/sec on a laptop with 8 GB VRAM and 32 GB RAM
- CPU-only server with 16 threads hits 18.4–21.0 tokens/sec at 78.8 GB memory
Read next
AI