PaperScope
LIVE · 2026-09-30 05:40 UTC

Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models

Yusheng Zhou, Eleanor Lin, David Jurgens

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36214 v1
Category
Submitted
2026-09-28

Abstract

Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.

Comment: 19 pages, 8 figures

arXiv abs page · PDF · same-day batch