PaperScope
LIVE · 2026-09-09 05:40 UTC

Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech

Chibuzor Okocha, Christan Grant, Zoey Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07968 v1
Category
Submitted
2026-09-07

Abstract

Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased

arXiv abs page · PDF · same-day batch