PaperScope
LIVE · 2026-09-29 05:40 UTC

AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35461 v1
Category
Submitted
2026-09-28

Abstract

As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.

Comment: Accepted to EMNLP 2026 Industry Track

arXiv abs page · PDF · same-day batch