PaperScope
LIVE · 2026-10-06 05:40 UTC

The ÌròyìnSpeech Text Corpus: 24,905 Curated Yorùbá Sentences for Speech and Language Technology

Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi, Iroro Orife, David Ifeoluwa Adelani

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05366 v1
Category
Submitted
2026-10-04

Abstract

ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.

Comment: 8 pages. Data descriptor for the ÌròyìnSpeech Text Corpus, doi:10.5281/zenodo.23138464. Under review at the Journal of Open Humanities Data

arXiv abs page · PDF · same-day batch