PaperScope
LIVE · 2026-09-09 05:40 UTC

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.08966 v1
Category
Submitted
2026-09-08

Abstract

Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.

arXiv abs page · PDF · same-day batch