WUSH-KV: 2-bit KV cache slashes memory bandwidth
WUSH-KV stores key/value elements in 2 bits while matching or beating other quantized methods in perplexity and staying close to full precision. A data-adaptive key transform and a value transform folded into model weights keep the cache in the 2-bit domain, cutting memory traffic proportionally from 16 bits.
- WUSH-KV stores keys and values in 2 bits instead of the usual 16
- It achieves the lowest perplexity among quantized transforms at every bitwidth
- The value transform is folded into model weights, keeping the cache in 2-bit
- A calibration phase is still needed to compute the key transform
Read next
AI