PaperScope
LIVE · 2026-09-11 05:40 UTC

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.11892 v1
Category
Submitted
2026-09-10

Abstract

As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an evaluation framework featuring diverse tasks and tailored metrics. Through this work, we aim to establish foundational infrastructures for Arabic Speech-LLMs under constraints imposed by limited Arabic speech resources.

arXiv abs page · PDF · same-day batch