PaperScope
LIVE · 2026-09-15 05:40 UTC

Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos

Feiyu Du, Xi He, Jia Li, Yapeng Tian, Weili Wu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14118 v1
Category
Submitted
2026-09-12

Abstract

Online understanding of who is talking to the camera wearer is a key capability for egocentric social interaction. However, existing talk-to-me (TTM) studies are commonly formulated as offline clip-level recognition, which is poorly aligned with online interaction and overlooks the diverse non-TTM speaking states that naturally arise in egocentric videos. In this paper, we revisit this problem by reformulating it as an online, frame-level prediction task. Instead of treating TTM as a binary problem against a single negative class, we model it in the presence of diverse and previously underexplored non-TTM states, such as talking-to-others, self-talking, and background conditions. To support this new formulation, we construct an Online TTM Dataset consisting of 406 egocentric video clips with approximately 900K annotated frames, each labeled with frame-level social interaction categories (e.g., background, TTM, talking-to-others, self-talking), by extending the Ego4D social interaction benchmark. In this benchmark, we evaluate five adapted baselines and develop a new model that integrates social cues across modalities. Experimental results show that our multimodal model, which jointly leverages audio, visual, and speech-semantic cues, achieves 75.5% frame-level F1 on TTM, outperforming strong baselines and enabling a systematic analysis of how different speaking states affect TTM recognition.

Comment: 10 pages, 4 figures, 4 tables. Accepted to the 34th ACM International Conference on Multimedia (MM '26)

arXiv abs page · PDF · same-day batch