| 1 | Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection | Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis | cs.AI | 2026-09-14 |
| 2 | Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering | Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah +1 | cs.CL | 2026-09-14 |
| 3 | HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses | Jieyuan Liu, Mengzhou Hu, Jefferson Chen +10 | cs.CL | 2026-09-14 |
| #4 | CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering | Sumit Barua, Guan Hong, Halil Dursunoglu +2 | cs.CL | 2026-09-14 |
| #5 | Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation | Yucheng Shen, Lingyong Yan, Jiulong Wu +4 | cs.AI | 2026-09-14 |
| #6 | Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models | Md Khalid Syfullah, Alvi Ataur Khalil | cs.CV | 2026-09-14 |
| #7 | CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense | Guo Fuzheng | cs.CR | 2026-09-14 |
| #8 | Where to Compute and How to Interact: Operator-Readable Adaptation with Gauge-Aware Transport | Zixuan Shen, Quanxu Wan, Bingchuan Wang +2 | cs.LG | 2026-09-14 |
| #9 | A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation | Yang Xing, Jiong Wu, Savas Ozdemir +11 | cs.CV | 2026-09-14 |
| #10 | Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA | Tristan Kirscher, Niklas C. Koser, Soren Pirk | cs.AI | 2026-09-14 |
| #11 | BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender | Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14 | cs.CV | 2026-09-14 |
| #12 | Long-to-Short Video Evidence Reasoning for Grounded Question Answering | Kaiyan Chen, Junbin Xiao, Xun Yang | cs.CV | 2026-09-14 |
| #13 | SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering | Davit Soselia, Joseph JaJa, Amitabh Varshney | cs.CV | 2026-09-14 |
| #14 | Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection | Zeyu Dong, Benjamin Wang, Joyee W. Jin | cs.CL | 2026-09-13 |
| #15 | Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return | Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi +3 | cs.SE | 2026-09-13 |
| #16 | Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation | Eduin E. Hernandez, Luis F. Garcia, Nurassyl Askar +2 | cs.CL | 2026-09-13 |
| #17 | Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing | Yakoub Bazi, Mohamad M. Al Rahhal, Mohamed A. Mekhtiche +1 | cs.CV | 2026-09-13 |
| #18 | NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass | Ali Derogar Odolou, Reza Nazari, Mostafa Salehi | cs.CL | 2026-09-13 |
| #19 | When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering | Saanvi Paturi, Arsen Kenzhebayev, Arham Sethi +3 | cs.CL | 2026-09-12 |
| #20 | T-SMART: Mechanism-Level Attribution for Tool-Augmented Time-Series Question Answering | Ivan Delgado, Himansi Gupta, Bishal Khatri +4 | cs.LG | 2026-09-12 |
| #21 | Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had | Achille Zappa | cs.AI | 2026-09-12 |
| #22 | GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning | Zhongyu Wang | cs.CL | 2026-09-12 |
| #23 | UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics | Yuzhe Li, Hao Yan, Hao Wang +6 | cs.AI | 2026-09-12 |
| #24 | Beyond OCR Accuracy: Text-Centric VQA Under Image Degradation with Modular and End-to-End | Ritali Vatsi, Rachapudi Jagadeesh, Shruti Singh Baghel +3 | cs.CV | 2026-09-12 |
| #25 | Realtime-Venus: A full-duplex interaction system with asynchronous delegation | Ruixiang Zhao, Hualei Wang, Renhe Sun +24 | cs.CV | 2026-09-12 |
| #26 | HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering | An Nguyen Phu, Dung Nguyen Quang, Luu Hieu An +3 | cs.CL | 2026-09-12 |
| #27 | Cost Characterization of Vertically Partitioned Federated Knowledge Graphs | Md Saikat Islam Khan Bappy, Oshani Seneviratne | cs.AI | 2026-09-12 |
| #28 | FedV-KGQA in Practice: Design Lessons and an Interactive Prototype | Md Saikat Islam Khan Bappy, Oshani Seneviratne | cs.AI | 2026-09-12 |
| #29 | Solar Intelligence | Jyotsna Singh | cs.AI | 2026-09-12 |
| #30 | Nuha-Speech: Building General-Purpose Arabic Speech-LLMs | Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi | cs.CL | 2026-09-10 |
| #31 | The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge | Jordi Luque, Lorenzo Concina, Marco Matassoni +2 | cs.CL | 2026-09-10 |
| #32 | ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs | Yizhan Li, Jianxin You, Mengyang Xiong +5 | cs.RO | 2026-09-09 |
| #33 | BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation | Karish Gupta, Matthew Alex, Alex Li +6 | cs.CV | 2026-09-09 |
| #34 | Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs | Ansuman Mullick, Eray Tüzün | cs.AI | 2026-09-09 |
| #35 | Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs | Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi | cs.CV | 2026-09-09 |
| #36 | What Should an Agent Forget? Separating What Is Stored from What Is Used | Yuhang Li, Yuchen Li | cs.AI | 2026-09-09 |
| #37 | LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation | Daniel Alejandro Coll Tejeda, Pedro García López, Daniel Barcelona-Pons | cs.IR | 2026-09-09 |
| #38 | The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs | Arquimedes Canedo | cs.CL | 2026-09-09 |
| #39 | Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning | Mingbo Yang, Wenqiang Wang, Zhaolu Kang +4 | cs.AI | 2026-09-09 |
| #40 | From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora | Christoph Wigbels, Ali Abusaleh, Markus T. Jansen +2 | cs.CL | 2026-09-09 |
| #41 | Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering | Zizhen Wang, Bo Feng, Zhengfeng Lai +5 | cs.CV | 2026-09-09 |
| #42 | Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators | Xinyu Chen, Adnan Mahmood, Mark Dras | cs.CV | 2026-09-09 |
| #43 | $S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants | Heyang Liu, Jiayi Huang, Wenyang Xiao +8 | cs.CL | 2026-09-09 |
| #44 | ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations | Jianjie Zheng, Peng Lai, Sijie Cheng +3 | cs.CL | 2026-09-09 |
| #45 | Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA | Yuexin Wu, Dayou Yu, Vasile Rus | cs.CL | 2026-09-09 |
| #46 | SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia | Jingyi Liao, Wenyu Zhang, Zhuohan Liu +6 | cs.CL | 2026-09-09 |
| #47 | ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance | Praphul Singh, Shanu Kumar, Akshat Agarwal +1 | cs.AI | 2026-09-08 |
| #48 | VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models | Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6 | cs.CV | 2026-09-08 |
| #49 | Evaluation of Contextual Understanding in Large Language Models | Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe +4 | cs.CL | 2026-09-08 |
| #50 | EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning | Jingpu Yang, Fengxian Ji, Mingxuan Cui +4 | cs.CV | 2026-09-08 |