PaperScope
LIVE · 2026-09-09 05:40 UTC

Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus

Lifeng Han, Jiahui Liang, Anna Latusek, Karim El Haff, Amal Haddad Haddad, Josua Höfgen, Kilian Evang, Min Ma, Maryia Zhyrko

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06634 v1
Category
Submitted
2026-09-06

Abstract

LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they are trained upon. To examine if Multiword Expressions (MWEs) still set a bottleneck for LLMs regarding language understanding and translation, we report the system performances from the WMT2026 Test Suites shared task, for which we used the publicly available multilingual parallel corpus AlphaMWE as the test suites. We received 31 MT systems' outputs covering English to Chinese (zh), Polish (pl), German (de), Arabic (ar) including Modern Standard Arabic (MSA) and two dialectal ones (Egyptian and Tunisian Arabic). We carried out automatic evaluations using BLEU, ChrF, BERT-score to select the Top3 systems per language pair, followed up with human evaluations on the selected systems. Our findings show that: figurative/MWE phenomena remain challenging; automatic metrics sometimes disagree; human evaluation uncovers language-specific errors hidden by aggregate scores.

Comment: WMT 2026 Test Suites Shared Task system paper (accepted). 20 pages, 8 figures, 7 tables

arXiv abs page · PDF · same-day batch