PaperScope
LIVE · 2026-10-09 05:40 UTC

FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs

Kyeong-Rae Kim, Sungnyun Kim, Tae-Hyun Oh

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.11310 v1
Category
Submitted
2026-10-08

Abstract

While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.

Comment: Project page: https://byulharang.github.io/FloorSAV/

arXiv abs page · PDF · same-day batch