PaperScope
LIVE · 2026-10-08 05:40 UTC

CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

Marcel Gibier, Thomas Thebaud, Olivier Boëffard, Jean-François Bonastre

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10208 v1
Category
Submitted
2026-10-07

Abstract

Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.

Comment: Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch