PaperScope
LIVE · 2026-09-10 05:40 UTC

Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion

Fadi Khatib, Meirav Galun, Ronen Basri

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.09491 v1
Category
Submitted
2026-09-08

Abstract

Camera pose estimation is a key step in 3D reconstruction and view-synthesis pipelines. We present a deep, global Structure-from-Motion framework based on learned view-graph aggregation. Our method employs a permutation-equivariant, edge-conditioned graph neural network that takes noisy pairwise relative poses as input and outputs globally consistent camera extrinsics. The network is trained without ground-truth supervision, relying solely on a relative-pose consistency objective. This is followed by 3D point triangulation and robust bundle adjustment. Our approach is efficient, scalable to more than a thousand images, and robust to graph density. We evaluate our method on MegaDepth, 1DSfM, Strecha, and BlendedMVS. These experiments demonstrate that our method achieves superior rotation and translation accuracy compared to deep track-centric methods while registering more images across many scenes, and competitive results compared to state-of-the-art classical pipelines, while being much faster.

Comment: Accepted to ECCV 2026. Project page: https://vgpa-sfm.github.io/

arXiv abs page · PDF · same-day batch