PaperScope
LIVE · 2026-09-29 05:40 UTC

Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech

Nicola Barbaro, Cristina Gena, Francesco Petriglia, Andrea Meirone, Alessandro Mazzei, Arianna Viotti

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33060 v1
Category
Submitted
2026-09-27

Abstract

Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-minute dataset of Italian child speech collected through a range of standardized diagnostic tests for children aged 3.5-6.5. Broca was optimized to handle phonetic variability in children's speech, including tone, accent, and speech errors, and achieved a state-of-the-art weighted Phoneme Error Rate of 13.36% on Italian speech. Remarkably, this performance was obtained using less than three hours of child-specific data, underscoring the model's efficiency and robustness in low-resource clinical settings. This work demonstrates that accurate, vocabulary-independent speech-to-IPA transcription can be achieved with minimal data, paving the way for more accessible, data-efficient tools to support speech assessment and diagnosis.

arXiv abs page · PDF · same-day batch