PaperScope
LIVE · 2026-09-29 05:40 UTC

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34223 v1
Category
Submitted
2026-09-28

Abstract

This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.

Comment: Preprint

arXiv abs page · PDF · same-day batch