| 1 | Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction | Kushal Patel, Pushkal Shrivastava, Mackenzie Lees +4 | cs.AI | 2026-09-14 |
| 2 | NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities | Haonan Jiang, Guojian Zhan, Jiancong Xie +6 | cs.AI | 2026-09-14 |
| 3 | Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models | Md Khalid Syfullah, Alvi Ataur Khalil | cs.CV | 2026-09-14 |
| #4 | ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs | Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui +9 | cs.CV | 2026-09-14 |
| #5 | A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation | Yang Xing, Jiong Wu, Savas Ozdemir +11 | cs.CV | 2026-09-14 |
| #6 | Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA | Tristan Kirscher, Niklas C. Koser, Soren Pirk | cs.AI | 2026-09-14 |
| #7 | AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation | Shengjie Jin, Zelong Sun, Hengbo Xu +2 | cs.CV | 2026-09-14 |
| #8 | Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs | Changxin Lu, Xiaoliang Meng, Yu Wu +5 | cs.RO | 2026-09-14 |
| #9 | Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders | Qingtao Xia, Jiahua Bao, Siyao Cheng +1 | cs.CV | 2026-09-14 |
| #10 | GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving | Xiao Liu, Haoyu Li, Jianghao Leng +2 | cs.CV | 2026-09-14 |
| #11 | Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing | Sehwan Park, Taehoon Kim, Geonhee Han +3 | cs.CV | 2026-09-13 |
| #12 | Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World | Guocun Wang, Kenkun Liu, Guorui Song +7 | cs.CV | 2026-09-13 |
| #13 | Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing | Yakoub Bazi, Mohamad M. Al Rahhal, Mohamed A. Mekhtiche +1 | cs.CV | 2026-09-13 |
| #14 | A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification | Leon Fernando, C Dombawala, P. Hettigoda +3 | cs.CV | 2026-09-13 |
| #15 | Two-Stage Mixture-of-LoRA for Multi-Task Medical Vision-Language Learning | Zhanghao Chen, Yuanyuan Li, Zhenyu Lu +3 | cs.CV | 2026-09-13 |
| #16 | E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning | Xiaoya Wang, Yutong Xu, Junjie Wang | cs.CL | 2026-09-13 |
| #17 | Vision-Language Models for Criterion-Level Grading of Handwritten Examinations in Outcome-Based Education | Md Khalid Syfullah, Asif Hasan Tonmoy, Saad Ahmed +1 | cs.CV | 2026-09-13 |
| #18 | SPARK: Representation-Level KV Memory Alignment for Safer Vision-Language Models | Mohd Azfar, Izhar Dad Khan | cs.CV | 2026-09-13 |
| #19 | RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes | Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra | cs.CV | 2026-09-12 |
| #20 | Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth | Tianhao Niu, Qingfu Zhu, Wanxiang Che | cs.CL | 2026-09-12 |
| #21 | Can Edge-Deployable Vision-Language Models Identify Species? | William Zhou, Mayukha Siripuram, Xiao Yan +2 | cs.AI | 2026-09-10 |
| #22 | ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding | Luca Della Libera, Cem Subakan, Mirco Ravanelli | cs.SD | 2026-09-10 |
| #23 | Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study | Yan Hon Michael Chung, Hanlin Wang | cs.LG | 2026-09-10 |
| #24 | Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models | Yixiang Liu, Zhongxing Xu, Zhonghua Wang +1 | cs.AI | 2026-09-10 |
| #25 | Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models | Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang +5 | cs.CV | 2026-09-10 |
| #26 | From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models | Meng Luo, Yicheng Liu, Jiahao Wang +5 | cs.CV | 2026-09-10 |
| #27 | Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions | Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden +3 | cs.CV | 2026-09-10 |
| #28 | Evaluation of Vision-Language Models Across Diverse Coastal Environments | Seth Knoop, Chad R. Samuelson, Gabriel R. Slade +2 | cs.CV | 2026-09-09 |
| #29 | BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation | Karish Gupta, Matthew Alex, Alex Li +6 | cs.CV | 2026-09-09 |
| #30 | MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery | Boshu Jia, Rongyu Chen, Linlin Yang +9 | cs.CV | 2026-09-09 |
| #31 | Show-Harness: Just a VLM Agent Can Play Robots | Yanzhe Chen, Zechen Bai, Zhijun Cao +7 | cs.RO | 2026-09-09 |
| #32 | Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization | Ayan Majumdar, Shounak Paul, Pushpdeep Singh +6 | cs.CL | 2026-09-09 |
| #33 | Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs | Xuan Cuong Ngo, Ngan Le | cs.CV | 2026-09-09 |
| #34 | Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection | Eli Schwartz | cs.CL | 2026-09-09 |
| #35 | From Pixels to Hierarchical Sequences: Quadtree Mask Encoding for Vision-Language Binary Change Detection | Xiao An, Ruikang Zhang, Chen Zhong +4 | cs.CV | 2026-09-09 |
| #36 | Vision-language models know more about agriculture than they show and rubric-grounded verifications close the gap | Earl Ranario, Jared Smith, Lars Lundqvist +1 | cs.CV | 2026-09-08 |
| #37 | VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models | Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6 | cs.CV | 2026-09-08 |
| #38 | Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs | Xiaofu Chen, Stella Frank, Yova Kementchedjhieva | cs.CV | 2026-09-08 |
| #39 | From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection | Mengzhe Geng, Yujia Lu, Patrick Littell +2 | cs.SD | 2026-09-08 |
| #40 | Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing | Xiaochuan Zhong, Yifan Hou, Chenxi Pang +1 | cs.CV | 2026-09-08 |
| #41 | CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation | Tinghe Ding, Jiahao Li, He Wang | cs.RO | 2026-09-08 |
| #42 | CLAMP: Constrained Decoding for Vision-Language Embodied Planning | Tianyi Ma, Parisa Kordjamshidi | cs.AI | 2026-09-08 |
| #43 | STSG-VQA: Evidence-Grounded Temporal Question Answering from Surgical Spatio-Temporal Scene Graphs | Jing Li, Duygu Sarikaya | cs.CV | 2026-09-08 |
| #44 | Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation | Shyam Nandan Rai, Francesco Di Salvo, Sebastian Doerrich +1 | cs.CV | 2026-09-08 |
| #45 | Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method | Boao Yu, Zimo Chen, Junreng Rao +4 | cs.CV | 2026-09-08 |
| #46 | CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs | Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7 | cs.CV | 2026-09-08 |
| #47 | CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning | SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2 | cs.CV | 2026-09-08 |
| #48 | 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints | Ziqin Huang, Yingyue Li, Chenyangguang Zhang +6 | cs.RO | 2026-09-08 |
| #49 | Drive by Hindsight and Foresight: Tool-Grounded Synergistic Reasoning over Hierarchical Memory for Autonomous Driving | Baojie Chen, Zijun Jia, Jing Zhong | cs.CV | 2026-09-08 |
| #50 | Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment | Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour +2 | cs.AI | 2026-09-08 |