Bonsai 2 27B compresses Qwen3.8 to 5.9GB while keeping 98.2% of quality
Prism ML released Ternary Bonsai 2 27B, a ternary-weight (1.76 bits per parameter) compression of Qwen3.8 27B that fits in a 5.9GB file. The company claims 98.2% of the original's average score, 262K-token context, image input and an Apache 2.0 license, while independent analysis flags limits in the benchmarks.
- 5.9GB file versus 54GB FP16 original — a 9.3x reduction
- Average score 83.9 versus 85.4 for the full model, math nearly unchanged
- RTX 5090 hits 129.9 tokens/s, beating H100 at 113.9 tokens/s
- Processing a full 262K-token context on M5 Pro takes about 11 minutes
Read next
AI