PaperScope
LIVE · 2026-09-23 05:40 UTC

A retrospective analysis on the use of LLMs to study infant syntax learning

Hélie Bazin, Anouk Barberousse, François Yvon

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.26539 v1
Category
Submitted
2026-09-22

Abstract

Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.

Journal: EMNLP 2026 Main Conference, ACL SIGDAT, Oct 2026, Budapest (Hungary), Hungary

arXiv abs page · PDF · same-day batch