PaperScope
LIVE · 2026-10-06 05:40 UTC

VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges

Sungjae Choi, Hanna Bae, Sunghyun Baek, Junmo Kim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06594 v1
Category
Submitted
2026-10-05

Abstract

Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.

Comment: Accepted to ACCV 2026

arXiv abs page · PDF · same-day batch