chiprook
← AI
AISeptember 25, 2026, 23:05

Moonshot AI releases Kimi Linear: hybrid attention cuts KV cache by 75%

Moonshot AI published a paper and open checkpoints for Kimi Linear, a 48B-parameter model (3B activated). Its hybrid architecture with the Kimi Delta Attention module reduces KV cache by 75% and speeds up decoding 6.3x at 1M-token context while matching full attention on benchmarks.

Moonshot AI releases Kimi Linear: hybrid attention cuts KV cache by 75%
#Moonshot#Kimi
Read next
AI

Grouped Value Attention cuts transformer KV cache by ~45%

AI

China's Moonshot AI puts Kimi K3 on Amazon Bedrock

AI

Chinese AI company Moonshot connects Kimi to Wall Street data providers

AI

Open Chinese Models Close Gap With Silicon Valley Frontier AI Models