PaperScope
LIVE · 2026-10-06 05:40 UTC

Learning What to Imitate: Entropy-Aware Distribution Mixing

Juan Garcia Giraldo, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin, Antoine Bosselut

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06671 v1
Category
Submitted
2026-10-05

Abstract

Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e., in a low-entropy state). To mitigate the degradation in generalisation and catastrophic forgetting caused by these conflicts, we propose \textbf{Entropy-Aware Mixing}: a dynamic per-token interpolation of the student and teacher distributions, gated by the student's predictive entropy. We implement both convex and geometric interpolations for both offline trace generation (via speculative decoding, then SFT) and on-policy forward-KL distillation. Our results show that entropy-aware mixing stabilises distillation, improving in-distribution and out-of-distribution math reasoning while better preserving general capabilities than fixed-teacher supervision. Nonetheless, the optimal entropy schedule depends on the training source, with offline-generated traces favouring concave schedules (greater overall teacher influence) and on-policy training favouring linear or convex schedules (teacher concentrated in high-entropy states).

Comment: 32 pages, 7 figures

arXiv abs page · PDF · same-day batch