180B Darwin model runs on a laptop without GPU via 4-bit GGUF
The POCKET-Darwin-180B-GGUF build packs the 180-billion-parameter MoE model Darwin-180B-RSI into 111 GB and runs it on CPU: a 16-thread server CPU hits 18.4–21 tokens/s, while an RTX 5060 laptop with 32 GB RAM reaches 4.17 tokens/s. On MMLU-Pro the quantized build matched the original at 87.65%.
- 4-bit build is 111 GB versus 360 GB in BF16
- Both versions scored 87.65% on 2,000 MMLU-Pro questions
- 16-thread server CPU reaches 18.4–21 tokens/s at 78.8 GB peak memory
- On SuperGPQA the quantized build scored 61.55% vs 59.10% for the base
Read next
AI