Google TurboQuant compresses KV cache to 3 bits with no accuracy loss
Google Research published TurboQuant, an algorithm that compresses the LLM KV cache to 3 bits per value with zero accuracy loss. For Llama-3.1-8B at 128K context, the cache shrinks from 16GB to about 2.6GB, with an 8x attention speedup on NVIDIA H100 GPUs.
- Llama-3.1-8B KV cache at 128K tokens drops from 16GB to ~2.6GB
- 4-bit TurboQuant delivered 8x attention speedup on NVIDIA H100
- Method combines PolarQuant with a 1-bit Johnson-Lindenstrauss transform
- Ported to llama.cpp, MLX, vLLM and SGLang within 24 hours
Read next
AI