| 1 | Prior-free relative 6D pose estimation of multiple object instances | Behdad Khodabandehloo, Andrea Caraffa, Davide Boscaini +1 | cs.CV | 2026-09-08 |
| 2 | Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views | Jinyu Miao, Jiusi Li, Yifei He +4 | cs.CV | 2026-09-08 |
| 3 | Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method | Boao Yu, Zimo Chen, Junreng Rao +4 | cs.CV | 2026-09-08 |
| #4 | CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs | Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7 | cs.CV | 2026-09-08 |
| #5 | 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints | Ziqin Huang, Yingyue Li, Chenyangguang Zhang +6 | cs.RO | 2026-09-08 |
| #6 | WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos | Giseong Hwang, Minjae Jo, Yeonghyeon Park +9 | cs.CV | 2026-09-08 |
| #7 | Zero-Shot 3D Plant Organ Segmentation with SAM3 and Semantic NeRFs | Andreas Gilson, Laura Hennig, Peter Pietrzyk | cs.CV | 2026-09-07 |
| #8 | I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models | Aydin Javadov, Daniel Schoess, Florian von Wangenheim | cs.LG | 2026-09-07 |
| #9 | Heat Kernel Textures: the Geodesic Gaussians That Do Not Splat | Simone Foti, Caner Korkmaz, Stefanos Zafeiriou +1 | cs.CV | 2026-09-07 |
| #10 | When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning | Chengliang Liu, Bo Li, Bob Zhang +3 | cs.CV | 2026-09-07 |
| #11 | Self-Supervised Multi-View 3D Gaze Target Estimation via Probabilistic Ray Marching | Keqi Chen, Vinkle Srivastav, Nicolas Padoy | cs.CV | 2026-09-07 |
| #12 | RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting | Hejun Wang, Jinxi Li, Junwei Jiang +4 | cs.CV | 2026-09-07 |
| #13 | KODAMA: Multimodal Digital Twin Reconstruction for Urban RF Propagation Modelling | Maximiliano Wardle, A. Ryo Koblitz | cs.CV | 2026-09-07 |
| #14 | MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling | Jin Xu, Xiaojian Huang, Zhuodong Luo +6 | cs.CV | 2026-09-07 |
| #15 | SkillAlign: Aligning Skill Interfaces for LLM-based Agents | Shuo Ren, Xiaomian Kang, Jiajun Zhang | cs.AI | 2026-09-07 |
| #16 | CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation | Jia-Jen Lee, Shih-Yen Hou, Kee Koon Ng +2 | cs.CV | 2026-09-07 |
| #17 | Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models | Mimo Shirasaka, Haochen Zhang, Yonatan Bisk | cs.CV | 2026-09-07 |
| #18 | DrugReason: Dynamic Multi-View Reasoning over Knowledge Graph and Language Evidence for Drug Repurposing | Zijie Liu, Hongxuan Li, Zhen Tan +5 | cs.LG | 2026-09-06 |
| #19 | 3DHarnessBench: Probing Agentic 3D-to-Code Capabilities of Frontier Vision-Language Models | Ling Liu, Bingchen Gong, Amal Dev Parakkat +1 | cs.CV | 2026-09-06 |
| #20 | Scaling 3D Generative Priors to Large-Scale Scene Meshes from Multi-View Images | SangEun Lee, Wonseok Chae, Hoyoung Yoo +3 | cs.CV | 2026-09-06 |
| #21 | DualPathOcc: Dual-Resolution BEV Encoder for 3D Occupancy Prediction | Lihao Qiu, Jian Chen, Ruihao Wang +3 | cs.CV | 2026-09-06 |
| #22 | WorldSculpt: Generating Compositional Worlds from Grounded Videos | Muyao Niu, Jixuan He, Ruihan Yu +9 | cs.CV | 2026-09-04 |
| #23 | CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation | Samer Abualhanud, Max Mehltretter | cs.CV | 2026-09-04 |
| #24 | Reflection-aware Generative Novel View Synthesis | GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh | cs.CV | 2026-09-04 |
| #25 | BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors | Vincent Leroy, Philippe Weinzaepfel, Lojze Zust +2 | cs.CV | 2026-09-04 |
| #26 | Learning 3D Editing without Paired Supervision via Generative Prior Distillation | Hao Wen, Weibin Yun, Hongxing Fan +4 | cs.CV | 2026-09-04 |
| #27 | An Evaluation Framework for Generating Multi-View Images of a Person in a Scene | Mahir Majid, Young Kyung Kim, Guillermo Sapiro | cs.CV | 2026-09-04 |
| #28 | BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data | Wei Zhang, Xin Li, Peishu Shi +4 | cs.CV | 2026-09-03 |
| #29 | Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving | Ruoyu Yao, Yusen Xie, Qingzhao Liu +5 | cs.CV | 2026-09-03 |
| #30 | Sparse auto-regressive modeling for scene generation from multi-view images | Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4 | cs.CV | 2026-09-03 |
| #31 | Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning | Yijun Yang, Shenghe Zheng, Wenbo Li +8 | cs.CV | 2026-09-03 |
| #32 | Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation | Haoyu Wang, Songchun Zhang, Haoran Li +3 | cs.CV | 2026-09-03 |
| #33 | TraveL: Transformer-based Multi-view Path Distributional Representation Learning | Fang He, Tao-yang Fu, Wang-chien Lee | cs.LG | 2026-09-03 |
| #34 | P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing | Yanshu Zhang, Shichong Peng, Mehran Aghabozorgi +2 | cs.CV | 2026-09-03 |
| #35 | MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion | Aidan Bradshaw, Marco Giordano, David Rode +8 | cs.CV | 2026-09-02 |
| #36 | InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling | Tianheng Lu, Guangyu Wang, Ruqi Huang +1 | cs.CV | 2026-09-02 |
| #37 | MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception | Guido Caccianiga, Sergey Prokudin, Yutong Chen +9 | cs.CV | 2026-09-02 |
| #38 | Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction | Zhiyuan Gao, Dominic Yurk, Yaser S. Abu-Mostafa | eess.IV | 2026-09-02 |
| #39 | TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views | Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9 | cs.CV | 2026-09-01 |
| #40 | Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation | R. James Cotton, Divya Joshi, Colleen Peyton | cs.CV | 2026-09-01 |
| #41 | iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy | Ravi Teja Vulchi, Carl Messerschmidt, Mohammadsadegh Vafaeinezhad +4 | cs.LG | 2026-09-01 |
| #42 | MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries | Utsab Ghosh, Roshni Chakraborty | cs.SD | 2026-09-01 |
| #43 | Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation | Yuanwang Yang, Buzhen Huang, Zongxuan Ren +2 | cs.CV | 2026-09-01 |
| #44 | Inverse Rendering for Modeling with Line Primitives | Kenji Tojo, Ariel Shamir, Nobuyuki Umetani +1 | cs.GR | 2026-09-01 |
| #45 | Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction | Xiaoyan Liu, Jiaxin Liu, Kangrui Li +1 | cs.CV | 2026-09-01 |
| #46 | Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration | Zekai Jin, Hanrong Zhang, Yihong Tang +3 | cs.RO | 2026-08-31 |
| #47 | FaceSnap: Real-Time Personalized Lightstage Facial Performance Capture | Rukhshanda Hussain, Noé Artru, Emeline Got +6 | cs.CV | 2026-08-31 |
| #48 | DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception | Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal +2 | cs.RO | 2026-08-31 |
| #49 | Multi-View Reflective Surface Inspection via Semantic-Saliency Cross-Verification | Van-Giang Nguyen, Thanh-Tuan Tran, Xuan-Hieu Phan +1 | cs.CV | 2026-08-31 |
| #50 | MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI | Athira J. Jacob, Puneet Sharma, Dorin Comaniciu +1 | cs.CV | 2026-08-31 |