PaperScope
LIVE · 2026-09-30 05:40 UTC

Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37852 v1
Category
Submitted
2026-09-29

Abstract

Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient's zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.

arXiv abs page · PDF · same-day batch