PaperScope
LIVE · 2026-10-02 05:40 UTC

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.01634 v1
Category
Submitted
2026-10-01

Abstract

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

Comment: 7 pages, 3 figures, 3 tables. Code and outputs: https://github.com/lazy-monster/yo-byt5

arXiv abs page · PDF · same-day batch