PaperScope
LIVE · 2026-09-09 05:40 UTC

Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities

Raphael Lorenzo-Louis, Bertrand Luvison, Serena Ivaldi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07394 v1
Category
Submitted
2026-09-07

Abstract

Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.

Comment: 5 pages, 4 figures. Workshop paper accepted to The 4th Workshop on Nonverbal Cues for Human-Robot Cooperative Intelligence at IROS 2026

arXiv abs page · PDF · same-day batch