PaperScope
LIVE · 2026-09-22 05:40 UTC

ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.22771 v1
Submitted
2026-09-19

Abstract

Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.

Comment: Accepted to Interspeech 2026. Project Website: https://nishitanand.github.io/paralinguistic-understanding-llm/

arXiv abs page · PDF · same-day batch