| 1 | Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language | Julian Truetsch, Felix Hauser, Christoph Stiller +1 | cs.CV | 2026-09-03 |
| 2 | Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection | Yue Zhao, Hua Yu, Yukun Zhao +6 | cs.CV | 2026-09-03 |
| 3 | When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection | Xuehao Wang, Jiaxin Hua, Runmei Li +4 | cs.CV | 2026-09-03 |
| #4 | Stereo 4D Radar for 3D Object Detection: Integrating Geometric Alignment and Absolute Velocity Estimation | Seung-Hyun Song, Dong-Hee Paek, Woong-Chan Byun +1 | cs.CV | 2026-09-02 |
| #5 | Information Density Imbalance in Visual Object Detection | Ziwei Zhao, Yanxi Lu, Yuwei Hu +8 | cs.CV | 2026-09-02 |
| #6 | Domain shift-robust object detection with GenAI image editing | Isabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel +4 | cs.CV | 2026-09-02 |
| #7 | If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection | Yinghao Sun, Shuguang Li, Jinliang Shao +1 | cs.CV | 2026-09-02 |
| #8 | CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer | Yisen Wang, Zhirong Wu, Limin Wang | cs.CV | 2026-09-01 |
| #9 | UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations | Chad Melton, PhD., Annabelle Kelton | cs.CV | 2026-09-01 |
| #10 | RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing | Kaiyue Kang, Qixuan He, Peijin Wang +9 | cs.CV | 2026-09-01 |
| #11 | Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving | Xin Zhou, Zongchuang Zhao, Zhibo Yang +13 | cs.CV | 2026-08-31 |
| #12 | Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling | Minghan Qin, Yuang Wang, Xiuyu Yang +6 | cs.CV | 2026-08-31 |
| #13 | A Composition-Aware Pretraining Framework for Geospatial Foundation Models | Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath +1 | cs.CV | 2026-08-31 |
| #14 | RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation | Quan Hao, Ziyang Tao, Chenxi Zhang +3 | cs.CV | 2026-08-31 |
| #15 | RailSyn: Diagnosis-Guided Image Generation for Traceable Data Completion in Railway Foreign Object Detection | Quan Hao, Chenxi Zhang, Ziyang Tao +6 | cs.CV | 2026-08-31 |
| #16 | Real-Time Scene-Adaptive Tone Mapping for High-Dynamic Range Object Detection | Gongzhe Li, Linwei Qiu, Peibei Cao +3 | cs.CV | 2026-08-31 |
| #17 | Seeing the Unseen: Camouflaged Object Detection Beyond the Visible Spectrum | Avi Gupta, Trasha Gupta | cs.CV | 2026-08-31 |
| #18 | A Lightweight Phenology-Aware YOLOv5 Framework for Tomato Growth Stage Detection in Resource-Constrained Bhutanese Greenhouse Environments | Sherab Gocha, Sou Nobukawa | cs.LG | 2026-08-30 |
| #19 | SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video | Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar | cs.CV | 2026-08-30 |
| #20 | SPLG-Mamba: Structure-Preserving Local-Global Mamba Network for Salient Object Detection in Optical Remote Sensing Images | Yi Xu, Ruichao Hou, Tongwei Ren +1 | cs.CV | 2026-08-30 |
| #21 | ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data Scarcity | Haoran Wang, Matthew Lau, Alec Helbling +7 | cs.CV | 2026-08-30 |
| #22 | Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs | Yu Cheng, Arushi Goel, Hakan Bilen | cs.CV | 2026-08-29 |
| #23 | RLG-TPV: Radar- and LiDAR-Guided Tri-Perspective View Fusion for Camera-Radar 3D Object Detection | Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates | cs.CV | 2026-08-29 |
| #24 | GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction | Yang Chen, Canyu Shen, Xinzhe Rao +5 | cs.CV | 2026-08-29 |
| #25 | Adversarial Calibration Attack on Autonomous Vehicles | Liangkai Liu, Qingzhao Zhang, Kang G. Shin | cs.RO | 2026-08-28 |
| #26 | Lossy Event Compression: From Event Stream Distortion to Task Performance | Zahra Rezaee, Catarina Brites, João Ascenso | cs.CV | 2026-08-28 |
| #27 | Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations | Marin Maletic, Marijana Peti, Tamara Petrovic +1 | cs.RO | 2026-08-28 |
| #28 | WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes | Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman +2 | cs.CV | 2026-08-28 |
| #29 | uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception | Trung Tien Dong, Zhenqi Wu, Aditya Penumarti +4 | cs.CV | 2026-08-28 |
| #30 | Variable-Granularity Tokenization for High-Resolution Object Detection | Khayrul Islam | cs.CV | 2026-08-27 |
| #31 | TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection | Su Wang, Yaochen Li, Min Yang +3 | cs.CV | 2026-08-27 |
| #32 | CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection | Hao Xu, Zhaoning Shi, Hehe Jin +1 | cs.CV | 2026-08-27 |
| #33 | Socialized Detector Learning: Trajectory-Guided and Reciprocal Distillation for Heterogeneous Object Detectors | Weihao Li, Yunqi Zhu, Zhihe Fan +5 | cs.CV | 2026-08-26 |
| #34 | TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection | Qiangqiang Zhou, Jiacong Yu, Jiawei Xu +3 | cs.CV | 2026-08-26 |
| #35 | MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection | Zhaojian Yao, Wei Gao, Tiesong Zhao +2 | cs.CV | 2026-08-26 |
| #36 | RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection | Zhuoyan Liu, Yihan Wang, Bo Wang +2 | cs.CV | 2026-08-26 |
| #37 | Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR | Sheethal Bhat, Bogdan Georgescu, Awais Mansoor +5 | cs.CV | 2026-08-25 |
| #38 | Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection | Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen +10 | cs.CV | 2026-08-25 |
| #39 | What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions | Yichao Gao, Yumo Zhang, Yunhao Yao +4 | cs.CR | 2026-08-25 |
| #40 | ROI-Gated SAHI: Content-Adaptive Slicing-Based Inference for Efficient Object Detection | Rashid Riyadh, Abd Ullah Khan, Imad Gohar +1 | cs.CV | 2026-08-25 |
| #41 | Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning | Simon Hakenes, Tobias Glasmachers | cs.LG | 2026-08-24 |
| #42 | Hyperbolic Hierarchical Clustering for Visual Representation Learning | Jianan Wei, Guikun Chen, Zhiyuan Weng +3 | cs.CV | 2026-08-24 |
| #43 | Cross-Generation Optimization of YOLOv26, YOLOv11, and YOLOv8 for Fine-Grained Small-Object Detection and Instance Segmentation in Complex Orchards | Ranjan Sapkota, Manoj Karkee | cs.CV | 2026-08-23 |
| #44 | DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection | Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin +2 | cs.CV | 2026-08-23 |
| #45 | Spiking Neural Networks for Energy-Efficient Object Detection in Forward-Looking Sonar Imagery | Gwenevere Frank, Gert Cauwenberghs | cs.CV | 2026-08-22 |
| #46 | C$^2$Path: Class-Conditional Pathway Decoupling for Vision-Language Incremental Object Detection | Lecheng Xu, Feifei Shao, Ouyangzi Ye +5 | cs.CV | 2026-08-22 |
| #47 | UniDiffFusion: A Unified Diffusion Framework for Multi-Task and Degradation-Robust Image Fusion | Xingxin Xu, Siqi Zhao, Xin Li +3 | cs.CV | 2026-08-22 |
| #48 | On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift | Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez +5 | cs.CV | 2026-08-21 |
| #49 | Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds | Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich +4 | cs.CV | 2026-08-21 |
| #50 | A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration | Jiekang Feng, Zhihe Fan, Yunqi Zhu +5 | cs.CV | 2026-08-21 |