PaperScope
LIVE · 2026-09-22 05:40 UTC

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Chenxu Xiong, Dongming Shen, Yuzhi Tang, Wentao Ma, Mu Li, Alex Smola

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.24812 v1
Category
Submitted
2026-09-21

Abstract

Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the Multi-Speaker Interaction Benchmark (MSI-Bench) for evaluating multi-speaker voice interaction. Each test case is a short multi-party multi-turn audio scene with participant context, expected tool calls, and atomic rubrics. The benchmark targets three capability families: multi-speaker memory, multi-speaker instruction following, and multi-speaker reasoning. It comprises 1,152 test cases, evenly split between Mandarin Chinese and English (576 each). The strongest configuration on each split passes all rubrics on only 66.8% of English and 54.5% of Mandarin cases, and the strongest open-weight configuration on 34.0% and 19.3%. Failure analysis separates perception from reasoning: open-weight models are bottlenecked by the multi-speaker audio front-end, while frontier systems still fail speaker-scoped decision making on clean transcripts---and models across the board often respond when no one has addressed them. These results identify speaker-grounded perception, speaker-scoped decision making, and conversational restraint as concrete targets for future voice agents.

Comment: 23 pages, 6 figures, 5 tables. Dataset: https://huggingface.co/datasets/M2cha4l1124/MSI-Bench ; Code: https://github.com/boson-ai/MSI-Bench

arXiv abs page · PDF · same-day batch