PaperScope
LIVE · 2026-10-09 05:40 UTC

LVSPM: Long Sequence View Synthesis and Pose Estimation Model

Xi Chen, Yachi Zhang, Linghao Chen, Minghua Liu, Hao Su, Zexiang Xu, Xiaoshuai Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10960 v1
Category
Submitted
2026-10-07

Abstract

We present LVSPM, a generalizable model that jointly estimates camera poses and synthesizes novel views from uncalibrated image collections. Trained with only RGB images and pose supervision, LVSPM avoids dense 3D ground truth and employs test-time training (TTT) layers to scale seamlessly to hundreds of input views. On RealEstate10k, Co3Dv2, and DL3DV, LVSPM surpasses VGGT in pose estimation across 16-256 views, with especially large margins at strict thresholds. For novel view synthesis under a practical protocol where more views cover larger scenes, LVSPM achieves state-of-the-art pose-free quality---surpassing even pose-dependent models in PSNR---and still maintains high quality as scene scale grows, while baselines collapse. The code is available at https://burningdust21.github.io/Projects/LVSPM .

Comment: ECCV 2026. Project Page: https://burningdust21.github.io/Projects/LVSPM/

arXiv abs page · PDF · same-day batch