PaperScope
LIVE · 2026-09-29 05:40 UTC

Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, Iheb Zouari, Essa Ahmed Abou Jabal, Ahmed Ezzat, Hind AL-Merekhi, Aisha Hamad M A Al-Naimi, Hadi Wazni, Bushra Alnajjar, Omar Amin, Haya Al-Thani, Houssam Eddine-Othman Lachemat, Marwa Elwakedy, Sundus Abdulmalik Al Nahari, Elahe Zahiri, Osamah Sarraj, Raghad Mousa, Mckeen Assi, Ahd Al Jumah, Heyam Salman, Alhanouf Abdulraqib, Sara Benoumhani, Alia Hamwi, Ayaat Al-Yasseri, Rim Ibrahim Ghazal, Lamia Ben hiba, Mohamed Eltabakh, Fatima Al-Raisi, Yassine El Kheir, Mohammed Abdulrahman, Hamdy Mubarak, Ayah Hashem, Lefkir Meriem, Ehsaneddin Asgari

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35564 v1
Category
Submitted
2026-09-28

Abstract

Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.

arXiv abs page · PDF · same-day batch