PaperScope
LIVE · 2026-10-06 05:40 UTC

Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training

Zhuojing Huang, Luise Pohlmann, Lisa Beinborn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06161 v1
Category
Submitted
2026-10-05

Abstract

During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.

Comment: EMNLP 2026, BabyLM Challenge; 18 pages, 6 figures

arXiv abs page · PDF · same-day batch