MIT and NVIDIA Close the Accuracy Gap in Native 8-bit LLM Training
Researchers from MIT and Carnegie Mellon identified and mathematically corrected the root cause of the residual accuracy gap in native FP8 LLM training. The Delta-Matching method restores the softmax Jacobian zero-row-sum invariant, enabling the 2x throughput of FP8 tensor cores without quality loss.
- On H100, FP8 reaches 3,958 teraflops versus 1,979 for BF16
- At 5.29B parameters, naive FP8 scored 16.3% on RULER-8K versus 53.5% for BF16
- Delta-Matching needs no architectural changes or smaller batch sizes
- The method reconciles E4M3 and E5M2 formats across the forward-backward boundary
Read next
AI