PaperScope
LIVE · 2026-10-05 05:40 UTC

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.03017 v1
Submitted
2026-10-02

Abstract

This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.

Comment: 8 pages, three figures, to be published in IEEE Speech Language Technology workshop 2026

arXiv abs page · PDF · same-day batch