PaperScope
LIVE · 2026-09-10 05:40 UTC

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.10153 v1
Category
Submitted
2026-09-09

Abstract

Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.

Comment: Accepted to EMNLP 2026

arXiv abs page · PDF · same-day batch