Grouped Value Attention cuts transformer KV cache by ~45%
Grouped Value Attention (GVA) reduces the persistent KV cache of transformers by roughly 45–47% with almost no accuracy loss: on a 350M-parameter model average accuracy differs from GQA by 0.01 points. DeepSeek-V4.1-Flash shrinks the global HBM footprint to 890 bytes per token and cuts its KV cache to about 1/8 of DeepSeek-V4-Flash.
- GVA cuts persistent KV cache by 45–47% versus GQA
- On a 350M model: 44.35 accuracy vs 44.36 for GQA and 43.88 for MLA
- DeepSeek-V4.1-Flash: 890 bytes of HBM per token, ~1/4 of predecessor
- SWA Bounded Replay shrinks KV cache to ~1/8 of DeepSeek-V4-Flash
Read next
AI