PaperScope
LIVE · 2026-10-01 05:40 UTC

Synthetic Data Characterization via Training Dynamics

Irene Lago, Ana Ezquerro, David Vilares

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39447 v1
Category
Submitted
2026-09-30

Abstract

Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.

Comment: Accepted at Findings of EMNLP 2026

arXiv abs page · PDF · same-day batch