chiprook
← AI
AISeptember 25, 2026, 12:00

Grouped Value Attention cuts transformer KV cache by ~45%

Grouped Value Attention (GVA) reduces the persistent KV cache of transformers by roughly 45–47% with almost no accuracy loss: on a 350M-parameter model average accuracy differs from GQA by 0.01 points. DeepSeek-V4.1-Flash shrinks the global HBM footprint to 890 bytes per token and cuts its KV cache to about 1/8 of DeepSeek-V4-Flash.

Grouped Value Attention cuts transformer KV cache by ~45%
#DeepSeek
Read next
AI

Chinese local governments subsidize AI filmmakers

AI

Microsoft Foundry adds native real-time voice agents to rival Gemini Live

AI

Fastino releases GLiNER2.5-Decide: a 340M open-weight decision model that runs on CPU

AI

Black Forest Labs releases FLUX 3 Action, a 7B robot control model