PaperScope
LIVE · 2026-09-29 05:40 UTC

Model-Aware Data Selection from In-and-Out Information Interplay

Yifan Wang, Xiaomin Li, Yuexing Hao, Dongwon Jung, Hemanth Neelgund Ramesh, Ananth Grama, Varun Chandrasekaran, Yu Hu, Andrzej Banburski-Fahey, Jaron Lanier

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33010 v1
Category
Submitted
2026-09-26

Abstract

LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and "know what they know." We observe an interesting rank equilibrium between knowledge stored in the weights and the data stream passing through the model. Across all model layers, we find that the hidden states (data stream) follow a U-shaped pattern, showing substantial compression in early layers and a steep rise during the late-layer decoding phase. In contrast, the weight rank follows an inverted U-shaped pattern, with very low rank in the early and late layers and high rank in the middle. We interpret this as an in-and-out information interplay: intermediate activations do not need to carry content that the weights can supply later, so they primarily preserve what the weights cannot provide. Motivated by this observation, we propose a model-aware data selection method, CAP (Counterfactual Assimilation Profile), which can determine whether a data candidate contains information accessible to the current model by utilizing the divergence gap in early- and late-layer representations between model-generated and reference responses. Across math, code, and science domains, CAP delivers 35.4% greater average improvement over the base model than the strongest baseline under different selection budgets. With only 10% of the data pool, CAP surpasses or matches full-pool training on math and science. We further show that CAP transfers to multimodal data selection and is robust to response horizon and noise.

arXiv abs page · PDF · same-day batch