PaperScope
LIVE · 2026-09-23 05:40 UTC

PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

Zhe Zhu, Yiheng Zhang, Peng Li, Zixing Zhao, Honghua Chen, Yaqing Zhang, Le Wan, Zhiyang Dou, Cheng Lin, Yuan Liu, Mingqiang Wei, Wenping Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.25832 v1
Category
Submitted
2026-09-22

Abstract

Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.

Comment: Accepted to SIGGRAPH Asia 2026 (ACM Transactions on Graphics). Project Page: https://czvvd.github.io/PartLLMPage/

arXiv abs page · PDF · same-day batch