PaperScope
LIVE · 2026-10-02 05:40 UTC

Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models

Zhanyu Chen, Jaap Jumelet

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.00540 v1
Category
Submitted
2026-09-30

Abstract

Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.

Comment: EMNLP Main 2026

arXiv abs page · PDF · same-day batch