REAL-Q uses dynamic gradient descent to fix LLM quantization
A new paper introduces REAL-Q, a post-training quantization method that replaces GPTQ's static Hessian with an aggregated Fisher matrix and dynamic block-wise gradient descent. It cuts KL divergence to 3.36 on LLaMA-3.1-8B versus 4.95 for GPTQ, with up to 49% reduction on Qwen3 models.
- KL divergence on LLaMA-3.1-8B: 3.36 for REAL-Q vs 4.95 for GPTQ
- Up to 49% lower KL divergence than GuidedQuant on Qwen3 models
- Adam step every 128 columns corrects accumulated quantization errors
- Works at W4A16 without retraining or a different inference engine
Read next
AI