Bonsai 2 27B: ternary weights shrink a 27B model to 5.9 GB
PrismML released Bonsai 2 27B on September 17, 2026, a Qwen3.8-27B-based model whose weights are reduced to three values (-1, 0, +1) using a blockwise Hadamard rotation. It occupies 5.9 GB in GGUF format versus roughly 54 GB in FP16 while retaining 98.2% of baseline benchmark performance.
- About 9x compression: 5.9 GB versus ~54 GB in FP16
- Average score 84.78 vs 86.32 for FP16 — 98.2% retention
- 1.72 bits per weight with g128 grouping and one shared FP16 scale
- Runs at ~47 tokens/s on Apple M5 Max and 91–130 tokens/s on RTX 4090
Read next
AI