PaperScope
LIVE · 2026-09-29 05:40 UTC

What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use

Yixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang, Yue Liu, Jun Zhang, Yuhong Liu, Jie Jiang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32991 v1
Category
Submitted
2026-09-26

Abstract

What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.

arXiv abs page · PDF · same-day batch