PaperScope
LIVE · 2026-10-06 05:40 UTC

Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

Ziyue WANG, T. Kanamori

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05126 v1
Category
Submitted
2026-10-04

Abstract

The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.

arXiv abs page · PDF · same-day batch