PaperScope
LIVE · 2026-09-03 05:40 UTC

Informational Antilocality and the Locality Bias in LLMs

Andrew McInnerney, Shane Storks, Steven Abney, Richard L. Lewis

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.27760 v1
Category
Submitted
2026-08-27

Abstract

We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.

arXiv abs page · PDF · same-day batch