| 1 | RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models | Canjie Liu, Jiawen Kang, Jinbo Wen +1 | cs.CV | 2026-09-02 |
| 2 | Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling | Ziquan Liu, Zhewei Zhu, Xuyang Shi | cs.CV | 2026-09-02 |
| 3 | ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering | Adrien Mialland, Marc Plantevit, Julien Gallois +1 | cs.IR | 2026-09-02 |
| #4 | TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models | Leqian Ding, Junning Qiu, Manwen Yang +2 | cs.CV | 2026-09-02 |
| #5 | LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory | Kun-Yang Yu, Yingzhe Li, Hongyu Xu +8 | cs.CV | 2026-09-02 |
| #6 | Towards Zero-Shot Transfer Across Embodiments For Driving VLAs | Caio Azevedo, Stefano Sabatini, Sascha Hornauer +1 | cs.CV | 2026-09-02 |
| #7 | YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification | Quansheng Hu, Qin Sun, Qiansen Dai +4 | cs.CV | 2026-09-02 |
| #8 | InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models | Chengyin Hu, Dingyi Lu, Jiaju Han +5 | cs.CV | 2026-09-02 |
| #9 | LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images | Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov | cs.CV | 2026-09-02 |
| #10 | TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval | Uicheol Jung, Juyoung Hong, Hojung Kwon +1 | cs.CV | 2026-09-02 |
| #11 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #12 | Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts | Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal +2 | cs.LG | 2026-09-02 |
| #13 | Test-Time Logit Prompting for Source-Free Missing Modality Adaptation | Taixi Chen, Nancy Guo | cs.CV | 2026-09-02 |
| #14 | Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification | Xuanbing Wen, Boxu Chen, Le Yang +4 | cs.CV | 2026-09-02 |
| #15 | Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework | Shuyao Xiao, Shengling Wang, Haoyu Niu +4 | cs.CV | 2026-09-02 |
| #16 | Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods | Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1 | cs.CV | 2026-09-01 |
| #17 | Video2Reaction: Training Foundation Video Models to Predict Audience Reaction | Sidong Zhang, Trang Nguyen, Shiv Shankar +4 | cs.CV | 2026-09-01 |
| #18 | DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation | Atif Belal, Lilian Hollard, Marco Pedersoli +1 | cs.CV | 2026-09-01 |
| #19 | MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models | Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5 | cs.CL | 2026-09-01 |
| #20 | AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $α$-Corrected Binary Cross Entropy and Factorized Latent Supervision | Jianzhong You, Yuan Gao, Chris McIntosh | cs.CV | 2026-09-01 |
| #21 | Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation | Haoyuan Deng, Haichao Liu, Wenkai Guo +6 | cs.RO | 2026-09-01 |
| #22 | Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers | Giovanni Bonetta, Matteo Merler, Davide Zago +2 | cs.AI | 2026-09-01 |
| #23 | FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making | Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza | cs.CV | 2026-09-01 |
| #24 | EdiTikZ: Scientific Figure Editing from Revision Trajectories | Christian Greisinger, Zhixue Zhao, Steffen Eger | cs.AI | 2026-09-01 |
| #25 | Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching | Jaewoo Park, Minyoung Lee, Sukmin Seo +11 | cs.RO | 2026-09-01 |
| #26 | IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals | Md. Atabuzzaman, Christian Alexander, Chris Thomas | cs.CV | 2026-09-01 |
| #27 | Reliability Challenges in Diffusion Vision-Language Models | Md. Atabuzzaman, Chris Thomas | cs.CV | 2026-09-01 |
| #28 | Agentic Multimodal Models for Environmental Hyperspectral Unmixing | Michał Cholewa, Luca Ciampi, Nicola Messina +2 | cs.CV | 2026-09-01 |
| #29 | EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents | Wei Wang, Wenqiao Zhang, Yutong Lin +14 | cs.RO | 2026-09-01 |
| #30 | REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs | Riyaaz Shaik, Chandru Venkataraman | cs.LG | 2026-09-01 |
| #31 | Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference | Reza Heidari, Hamed R. Tavakoli, Juho Kannala | cs.CV | 2026-09-01 |
| #32 | Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models | Jiayu Ding, Zhuodong Liu, Lei Zhang +6 | cs.CV | 2026-09-01 |
| #33 | From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding | Raul Ortega, José Manuel Gómez-Pérez | cs.CV | 2026-09-01 |
| #34 | A multicenter benchmark and clinically structured metric for coronary CTA report generation | Zhiyu Ye, Yue Sun, Limiao Zou +7 | cs.CV | 2026-09-01 |
| #35 | Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting | Udo Schlegel, Shubhangi, Gabriel Dax +3 | cs.CV | 2026-09-01 |
| #36 | Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning | Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee +2 | cs.AI | 2026-09-01 |
| #37 | The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence | Genpei Zhang | cs.CV | 2026-09-01 |
| #38 | Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation | Yumi Lee, Harim Oh, Hyoryung Kim +52 | cs.CV | 2026-09-01 |
| #39 | Towards Generalizable Visually Grounded Exploration of Household Devices | Linhao Zheng, Zeming Liu, Wangke Chen +4 | cs.AI | 2026-09-01 |
| #40 | Visual Attention Faithfulness in Vision-Language Models is Heterogeneous | Xurui Song, Weishi Wang, Zhongqi Yue +5 | cs.CV | 2026-09-01 |
| #41 | Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis | Minsik Choi, Geewook Kim, Young Geun Kim | cs.LG | 2026-09-01 |
| #42 | EarthLD: Towards Unified Open-World Landslide Understanding via Vision-Language Guided Diffusion Models | Yuanchao Su, Lianru Gao, Mengying Jiang +3 | cs.CV | 2026-09-01 |
| #43 | Controllable Image Captioning with Prompt-Conditioned Scene Rewards | Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim | cs.CV | 2026-09-01 |
| #44 | From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers | Siyi Liu, Hanjun Yang, Chenchen Zhang +7 | cs.IR | 2026-09-01 |
| #45 | Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures | Can Polat, Mustafa Kurban, Erchin Serpedin +1 | cs.CV | 2026-09-01 |
| #46 | Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning | Kaizhen Tan, Yang Feng, Heqing Du +3 | cs.CV | 2026-09-01 |
| #47 | You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change | Kaizhen Tan | cs.CV | 2026-09-01 |
| #48 | ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs | Yuta Kato, Shintaro Ozaki, Kazuki Hayashi +4 | cs.CL | 2026-09-01 |
| #49 | Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation | Teresa DiMeola, Charles Walter, Hong Xiao | cs.CV | 2026-09-01 |
| #50 | BrainDiff: Longitudinal Report Generation for Multimodal Brain MRI | Krish Patel, Peirong Liu | cs.CV | 2026-09-01 |