PaperScope
LIVE · 2026-09-11 05:40 UTC

SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views

Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.11279 v1
Category
Submitted
2026-09-10

Abstract

With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target instance, jointly improving shape accuracy and achieving object-level disentanglement without multi-stage pipelines. To ensure reconstruction stability, a lightweight Spatial RankGNN selects the optimal reference view with a selection accuracy of 73.5\%. Extensive experiments demonstrate that our method boosts average reconstruction precision by 11\% across various metrics compared to state-of-the-art baselines. These results reveal a strong instance-disentanglement capability and clear benefits for driving, robotics, AR/VR, and heritage digitisation.

arXiv abs page · PDF · same-day batch