PaperScope
LIVE · 2026-09-22 05:40 UTC

Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs

Mai Mohamed Eida, Ryan Dolan, Paul de Nijs, Jonathan Dunn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23955 v1
Category
Submitted
2026-09-21

Abstract

Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Arabic (SEA) dialect both in LLM and resource development. Does this lack of representation influence an LLM's view of the acceptability of SEA (upstream), and does an upstream bias against SEA lead to worse performance (downstream)? We investigate the upstream effect of SEA dialectal features on LLM preferences in a Targeted Syntactic Evaluation (TSE) task which reveals a significant bias against SEA across multiple LLMs. We then analyze the effect of these same features on downstream model performance on MMLU benchmarks and show that models experience a degradation in performance when presented with SEA. This work highlights the need for further exploration on how sub-dialectal variation impacts language technologies.

arXiv abs page · PDF · same-day batch