PaperScope
LIVE · 2026-09-23 05:40 UTC

Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models

Hongjin Song, Runwu Shi, Weiqiao Shan, Jiale Luo, Yujin Wang, Yifei Wu, Chunxiang Jin

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.25890 v1
Category
Submitted
2026-09-22

Abstract

Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention, and token weights under batch-mean loss. We disentangle these factors through matched comparisons. In the tested settings, short-to-long ordering shows no independent benefit when batch composition and token exposure are fixed. First-epoch grouping lowers perplexity for Mimi under batch-mean loss, but this gain is not observed under token-balanced loss. The cross-tokenizer results are consistent with a link between chunk-length variation and token weighting. This work provides a systematic analysis protocol for studying length-based training in variable-length speech models.

arXiv abs page · PDF · same-day batch