chiprook
← AI
AISeptember 30, 2026, 23:05

REAL-Q uses dynamic gradient descent to fix LLM quantization

A new paper introduces REAL-Q, a post-training quantization method that replaces GPTQ's static Hessian with an aggregated Fisher matrix and dynamic block-wise gradient descent. It cuts KL divergence to 3.36 on LLaMA-3.1-8B versus 4.95 for GPTQ, with up to 49% reduction on Qwen3 models.

REAL-Q uses dynamic gradient descent to fix LLM quantization
#LLaMA#Qwen
Read next
AI

Chat template triggers "I'm just an AI" disclaimer in 8 LLMs

AI

Nokia open-sources AnyJev: a training-free layer that turns any open LLM into a calibrated decision model

AI

PrismML hopes its tiny LLM could change how we all use AI

AI

Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight