PaperScope
LIVE · 2026-09-09 05:40 UTC

Physico-Geospatial Grounded Scene Interpretation for Mobile Robotics

Nicolas Schuler, Janik Kurtz, Lea Dewald, Marcel Sauber, Félicia Teferle, Jürgen Graf

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06629 v1
Category
Submitted
2026-09-06

Abstract

Recent advancements in deep learning allow robotic agents to interact with dynamic and unstructured environments. Of special interest is the integration of physico-geospatial world knowledge into such systems, either by using physics-aware machine learning models, knowledge graphs to model relationships or spatio-temporal and logical reasoning. In the present work, we introduce an approach to augment the output of pre-trained, unmodified VLMs used for scene interpretation by integrating semantic descriptions, OpenStreetMap building data and street information with positional, temporal and metric information obtained from our sensory systems, fusing this information using LLMs. We apply this concept to an outdoor recording within a university campus, achieving an F1-Score of 0.83 in the task of grounding buildings and 0.64 for path surface grounding on our pilot evaluation set. The results demonstrate the conceptual capability of the proposed solution to deliver physico-geospatial grounded natural language descriptions. Code and results are available at https://datahub.rz.rptu.de/hstr-csrl-public/publications/physico-geospatial-grounded-scene-interpretation

Comment: 9 pages, 3 figures, 3 table; accepted for International Conference on FutureTech 2026 (ICFT), AMMAN, JORDAN October 18-22 2026

arXiv abs page · PDF · same-day batch