PaperScope
LIVE · 2026-09-23 05:40 UTC

ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

Pu Wang, Yujun Wang, Hugo Van hamme

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.26648 v1
Submitted
2026-09-22

Abstract

Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.

Comment: Submitted to IEEE ICASSP 2027

arXiv abs page · PDF · same-day batch