Moonshot AI releases Kimi Linear: hybrid attention cuts KV cache by 75%
Moonshot AI published a paper and open checkpoints for Kimi Linear, a 48B-parameter model (3B activated). Its hybrid architecture with the Kimi Delta Attention module reduces KV cache by 75% and speeds up decoding 6.3x at 1M-token context while matching full attention on benchmarks.
- Model: 48B total parameters, 3B activated per token (MoE)
- 75% smaller KV cache, 6.3x faster decoding at 1M tokens
- RULER at 1M context: 94.8; at 128k: 84.3 with 3.98x speedup
- Flagship Kimi K3 uses 69 KDA layers and 24 Gated MLA layers
Read next
AI