PaperScope
LIVE · 2026-09-29 05:40 UTC

Muon Sublates the Edge of Stability in LLM Pretraining

Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu, Fanghui Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34915 v1
Category
Submitted
2026-09-28

Abstract

Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2ρ_b/η$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining

Comment: 30 pages, 19 figures

arXiv abs page · PDF · same-day batch