| 1 | RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation | Xiaolei Lang, Ze Kang, Zehao Huang +1 | cs.CV | 2026-09-02 |
| 2 | InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling | Tianheng Lu, Guangyu Wang, Ruqi Huang +1 | cs.CV | 2026-09-02 |
| 3 | Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations | Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran | cs.CV | 2026-09-02 |
| #4 | AffectDelta: Beyond Emotion Labels for Image Editing | Xingzu Zhan, Lin Gu, Ruogu Fang | cs.CV | 2026-09-02 |
| #5 | MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval? | Uicheol Jung, Juyoung Hong, Geuntaek Lim +1 | cs.CV | 2026-09-02 |
| #6 | WiFlow: Estimating Optical Flow using WiFi Channel State Information | Thomas Weigel, Simon Kiefhaber, Fabian Portner +2 | cs.CV | 2026-09-02 |
| #7 | Towards Zero-Shot Transfer Across Embodiments For Driving VLAs | Caio Azevedo, Stefano Sabatini, Sascha Hornauer +1 | cs.CV | 2026-09-02 |
| #8 | If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection | Yinghao Sun, Shuguang Li, Jinliang Shao +1 | cs.CV | 2026-09-02 |
| #9 | CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation | Shucheng Zhang, Yuang Zhang, Bingzhang Wang +3 | cs.RO | 2026-09-02 |
| #10 | T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation | Yan Wang, Xinyi Hou, Weiguo Lin +2 | cs.CV | 2026-09-02 |
| #11 | DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space | Steffen Hagedorn, Aron Distelzweig, Alexandru P. Condurache | cs.RO | 2026-09-02 |
| #12 | CC-4DGS: Computational Deformation and Point-Cloud Compression for Storage-Efficient Dynamic Gaussian Splatting | Kyungdae Park, Chae Eun Rhee | cs.CV | 2026-09-02 |
| #13 | World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models | Chuhan Zhang, Seiji Ito, Kenta Hoshino +2 | cs.CV | 2026-09-02 |
| #14 | KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection | Zhouyuan Huai, Meiqi Wan, Yan Yang +4 | cs.CV | 2026-09-02 |
| #15 | Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation | Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote | cs.RO | 2026-09-02 |
| #16 | GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric Matching | Lu Han, Xiting Sun, Hao Wang +4 | cs.CV | 2026-09-02 |
| #17 | TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views | Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9 | cs.CV | 2026-09-01 |
| #18 | Consistency as Regularization for Unsupervised Shadow Removal | Anh-Kiet Duong, Petra Gomez-Krämer, Jean-Michel Carozza | cs.CV | 2026-09-01 |
| #19 | Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation | Ehsan Faghih, Fatemeh Ashrafi, Marguerite Moore +1 | cs.CV | 2026-09-01 |
| #20 | Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison | Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit | cs.CV | 2026-09-01 |
| #21 | TempCloze: Can Video-LLMs Identify the Missing Middle? | Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4 | cs.CV | 2026-09-01 |
| #22 | Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds | Clinton Enwerem, John S. Baras, Calin Belta | cs.RO | 2026-09-01 |
| #23 | Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept Erasure | Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang +2 | cs.CV | 2026-09-01 |
| #24 | MeshSplatBench: A Unified Benchmark for Triangle-Based Neural Rendering | Kaixuan Zhang, Minxian Li, Mingwu Ren +1 | cs.GR | 2026-09-01 |
| #25 | Agentic Multimodal Models for Environmental Hyperspectral Unmixing | Michał Cholewa, Luca Ciampi, Nicola Messina +2 | cs.CV | 2026-09-01 |
| #26 | Seeing the World and the Self from Egocentric Video | Kai Guan, Minchao Jiang, Ruichen WangLi +2 | cs.CV | 2026-09-01 |
| #27 | MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation | Zhijian Qiao, Xinjiang Wang, Jiajie Chen +5 | cs.CV | 2026-09-01 |
| #28 | EvoGS: Modeling Deformation Evolution for Dynamic Gaussian Splatting | Wei Dong, Shahram Shirani, Jun Chen +1 | cs.CV | 2026-09-01 |
| #29 | PredErase: Training-Free Object-and-Effect Removal with Predictive Latent Guidance | Waikit Xiu, Qiang Lu, Junbiao Chen +1 | cs.CV | 2026-09-01 |
| #30 | CERF: Communication-Efficient and Retraining-Free Collaborative Perception | Jiuwu Hao, Ziyi Ni, Liguo Sun +5 | cs.CV | 2026-09-01 |
| #31 | From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding | Raul Ortega, José Manuel Gómez-Pérez | cs.CV | 2026-09-01 |
| #32 | Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking | Orcun Cetintas, Guillem Brasó, Tim Meinhardt +1 | cs.CV | 2026-09-01 |
| #33 | On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios | Zhe Shen, Liyuan Lou, Yifei Yu +4 | cs.CV | 2026-09-01 |
| #34 | HELIOS: From midnight to noon, continuous outdoor urban scene relighting | Hala Djeghim, Nathan Piasco, Luis Roldão +4 | cs.CV | 2026-09-01 |
| #35 | Can Scene Text Recognition Read Rare Compositions? | Genpei Zhang | cs.CV | 2026-09-01 |
| #36 | RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing | Kaiyue Kang, Qixuan He, Peijin Wang +9 | cs.CV | 2026-09-01 |
| #37 | VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM | Sangmin Song, Sarath Kodagoda, Marc G. Carmichael +4 | cs.CV | 2026-09-01 |
| #38 | Controllable Image Captioning with Prompt-Conditioned Scene Rewards | Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim | cs.CV | 2026-09-01 |
| #39 | Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning | Zixuan Wang, Yixin Hu, Wen Li +4 | cs.CV | 2026-09-01 |
| #40 | You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change | Kaizhen Tan | cs.CV | 2026-09-01 |
| #41 | Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation | Teresa DiMeola, Charles Walter, Hong Xiao | cs.CV | 2026-09-01 |
| #42 | BeamRMX: Radiation-Pattern-Driven Learning for Generalizable Beam Radio Map Prediction and Beam Management | Yue Zhang, Xiucheng Wang, Wenshuo Chen +1 | eess.SP | 2026-09-01 |
| #43 | EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection | Guanzhong Sun, Junyi Ma, Yuxuan Wu +1 | cs.LG | 2026-09-01 |
| #44 | Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration | Zekai Jin, Hanrong Zhang, Yihong Tang +3 | cs.RO | 2026-08-31 |
| #45 | Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation | Vida Adeli, Soroush Mehraban, Jacob Rommann +3 | cs.CV | 2026-08-31 |
| #46 | CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction | Zhengxu Tang, Guofeng Cui, Ziyu Gong +8 | cs.CV | 2026-08-31 |
| #47 | Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving | Xin Zhou, Zongchuang Zhao, Zhibo Yang +13 | cs.CV | 2026-08-31 |
| #48 | SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies | Weiqi Wang, Zhi Li, Yudong Lei +7 | cs.RO | 2026-08-31 |
| #49 | BRF-GS: Hyperspectral Bidirectional Reflectance Factor Modeling and Image Generation Based on 3D Gaussian Splatting | Yiling Yao, Wenjuan Zhang, Bowen Wang +3 | cs.CV | 2026-08-31 |
| #50 | Driving on Memory | Christian Löwens, Thorben Funke, Alexandru Paul Condurache | cs.CV | 2026-08-31 |