Hybrid-precision attention reduces compute cost with minimal accuracy loss
The HyQuant method quantizes most query, key, and value tensors to low precision, keeping only critical tokens and a local sliding window in full precision. This yields 1.32–3.58x speedup of the decoding kernel and 1.04–1.17x end-to-end decoding with less than 1% accuracy loss.
- Decoding kernel speedup 1.32–3.58x, end-to-end decoding 1.04–1.17x
- Accuracy loss less than 1% on long-context and reasoning benchmarks
- Overhead of critical token detector is 3–5% of runtime
- Triton kernel patch ready, pluggable with one import
Read next
AI