| 1 | SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators | Yuncong Yang, Zhengtao Han, Furkan Ozyurt +6 | cs.CV | 2026-09-08 |
| 2 | Point4D: Long-range 4D Motion Reconstruction | Minsik Jeon, Jay Karhade, Deva Ramanan +1 | cs.CV | 2026-09-08 |
| 3 | Studying Image Tokenizers as Visual Languages in Unified Multimodal Models | Siting Li, Zhengyang Wang, Simon Shaolei Du +2 | cs.CV | 2026-09-08 |
| #4 | Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs | Xiaofu Chen, Stella Frank, Yova Kementchedjhieva | cs.CV | 2026-09-08 |
| #5 | Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout | Zhuoran Zhao, Shengju Qian, Tongtong Liang +7 | cs.CV | 2026-09-08 |
| #6 | DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination | Yankai Fu, Ning Chen, Junkai Zhao +5 | cs.RO | 2026-09-08 |
| #7 | "World Knowledge" in the Weights: Reading Concept Circuits of Vision Transformers | Yanlin Chen, Tang Li, Xi Peng | cs.CV | 2026-09-08 |
| #8 | Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild | Fei Teng, Sheng Wu, Mengfei Duan +7 | cs.CV | 2026-09-08 |
| #9 | EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning | Jingpu Yang, Fengxian Ji, Mingxuan Cui +4 | cs.CV | 2026-09-08 |
| #10 | DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding | Zhuo Cao, Bingqing Zhang, Sen Wang +1 | cs.CV | 2026-09-08 |
| #11 | Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views | Jinyu Miao, Jiusi Li, Yifei He +4 | cs.CV | 2026-09-08 |
| #12 | Evaluation Principles for MRI-MRA Registration in Trigeminal Neuralgia: An ROI-Centered Neurovascular Benchmark | Xupeng Zhang, Xihang Wang, Michael Xie +6 | cs.CV | 2026-09-08 |
| #13 | Compensating for Scarce Historical Images in Cross-Domain Cultural Heritage Retrieval Using Synthetic Aging | Marcin Iwanowski, Adam Mazgaj, Ferdynand Gorski +1 | cs.CV | 2026-09-08 |
| #14 | Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics | Ruibo Ming, Lei Sun, Deheng Zhang +8 | cs.CV | 2026-09-08 |
| #15 | CVT-GS: Learning to Simplify 3D Gaussian Splatting with Centroidal Voronoi Tessellation | Bingxian Li, Yilong Li, Jingliang Peng +6 | cs.CV | 2026-09-08 |
| #16 | MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis | Hanyi Zhang, Maximilian Hoermann, Lion J. Gleiter +5 | cs.MA | 2026-09-08 |
| #17 | Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing | Xiaochuan Zhong, Yifan Hou, Chenxi Pang +1 | cs.CV | 2026-09-08 |
| #18 | From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video | Qiaohui Chu, Haoyu Zhang, Meng Liu +3 | cs.CV | 2026-09-08 |
| #19 | MFVINS: Multiple Fisheye Camera-Based Visual Inertial System | Eunseong Jang, YuJin Chung, Sang Jun Lee +2 | cs.RO | 2026-09-08 |
| #20 | CLAMP: Constrained Decoding for Vision-Language Embodied Planning | Tianyi Ma, Parisa Kordjamshidi | cs.AI | 2026-09-08 |
| #21 | Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance | Luyao Tang, Bingjun Luo, Dong Yi +5 | cs.CV | 2026-09-08 |
| #22 | SignRefine: Adapting Foundational Video Models for Sign Language Generation | Anton Pelykh, Edward Fish, Ozge Mercanoglu Sincan +1 | cs.CV | 2026-09-08 |
| #23 | AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction | Feiyu Zhao, Yuetong Li, Chenxi Xiao | cs.RO | 2026-09-08 |
| #24 | AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation | Shanwei Fan, Bin Zhang, Zhiwei Xu +4 | cs.CV | 2026-09-08 |
| #25 | Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method | Boao Yu, Zimo Chen, Junreng Rao +4 | cs.CV | 2026-09-08 |
| #26 | From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs | Juwan Chung, Sungjune Park, Yeongyun Kim +1 | cs.CV | 2026-09-08 |
| #27 | Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models | Bella Godiva, Yeonju Kim, Yong Man Ro | cs.SD | 2026-09-08 |
| #28 | Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking | Jue Wang, Xuan Wang, Hao Zhou +6 | cs.CV | 2026-09-08 |
| #29 | CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs | Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7 | cs.CV | 2026-09-08 |
| #30 | RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation | Jingxuan Zhu, Jingyi Li, LiangLiang Chen +3 | cs.RO | 2026-09-08 |
| #31 | Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs? | Bangshuo Zhu, Wei Song, Yuxin Cao +4 | cs.CV | 2026-09-08 |
| #32 | EMBLEM: Enhancing Multi-script Table Detection through Masking | Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan | cs.LG | 2026-09-08 |
| #33 | Supervised Cross-Modal Feature Alignment for Zero-Wearable Freezing of Gait Detection in Parkinsonism | Aryan Singh, Chandan Biswas | cs.CV | 2026-09-08 |
| #34 | From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment | Aoting Zhang, Mingze Gao, Dongbao Yang +5 | cs.CV | 2026-09-08 |
| #35 | Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models | Ke Hao, Yuanzhi Liang, Tingxi Chen +5 | cs.CV | 2026-09-08 |
| #36 | ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing | Yi Ting Shen, Kentaroh Toyoda, Alex Leung | cs.CR | 2026-09-08 |
| #37 | CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations | Aman Mehta, Riya Baviskar | cs.RO | 2026-09-08 |
| #38 | CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning | SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2 | cs.CV | 2026-09-08 |
| #39 | SoftRerank: Hierarchical Soft Fusion with Candidate-Label Reranking for Long-Tailed Micro-Action Recognition | Yichi Zhang, Zhichao Xia, Yanjun Chi +6 | cs.CV | 2026-09-08 |
| #40 | PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation | Cong Wang, Hanxin Zhu, Yonglin Tian +5 | cs.CV | 2026-09-08 |
| #41 | Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation | Jianqiang Xiao, Xiang Deng, Yuexuan Sun +3 | cs.RO | 2026-09-08 |
| #42 | SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code | Wentao Li, Yibo Wu, Yizhe Chen +5 | cs.CV | 2026-09-08 |
| #43 | SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation | Soroush Mehraban, Xin Lei Lin, Vida Adeli +5 | cs.CV | 2026-09-08 |
| #44 | RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts | Diwas Lamsal, Juha Carlon, Reinhard Claeys +8 | cs.AI | 2026-09-08 |
| #45 | VI-Bench: Benchmarking Prompt Inversion from AIGC Videos | Wulin Xie, Rui Zhao, Kecen Li +5 | cs.CV | 2026-09-08 |
| #46 | BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset | Md. Sadman Sakib, Zisan Mahmud, Md. Fahim Arefin +1 | cs.CL | 2026-09-07 |
| #47 | Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning | Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2 | cs.AI | 2026-09-07 |
| #48 | A Black-Box Adversarial Attack on Human Pose Estimation and Keypoint-Based Action Recognition Models | Kacper Mroczek, Michal Kepski | cs.CV | 2026-09-07 |
| #49 | TaskGuard: Task-Conditioned Restoration Utility for Risk-Aware Object Detection | Vung Pham | cs.CV | 2026-09-07 |
| #50 | ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding | Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang +2 | cs.CV | 2026-09-07 |