| 1 | Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework | Cagri Temel | cs.RO | 2026-09-02 |
| 2 | Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems | Yihang Chen, Yuxiang Chen, Yuxuan Huang +3 | cs.AI | 2026-09-02 |
| 3 | Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills | Jianlyu Chen, Yuyang Hu, Hongjin Qian +8 | cs.AI | 2026-09-02 |
| #4 | Dimension Dependent Correlation Gap Bounds under Restricted Independence | Arjun Ramachandra | math.PR | 2026-09-02 |
| #5 | TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis | Jiahao Wu, Zhenqun Yang, Chen Jason Zhang +1 | cs.LG | 2026-09-02 |
| #6 | Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework | Yan Zhong, Gefei Chen, Qiufang Ma +4 | cs.CV | 2026-09-02 |
| #7 | When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models | Smitha Muthya Sudheendra, Jaideep Srivastava | cs.CL | 2026-09-02 |
| #8 | Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment | Chenyu Zhou, Qiliang Jiang, Shuning Wu +1 | cs.LG | 2026-09-02 |
| #9 | YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification | Quansheng Hu, Qin Sun, Qiansen Dai +4 | cs.CV | 2026-09-02 |
| #10 | Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization | Bizhe Bai, Jiakang Yuan, Hongming Wu +8 | cs.CL | 2026-09-02 |
| #11 | Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics | Yijie Lin, Ching-Chun Chang, Isao Echizen +2 | cs.CR | 2026-09-02 |
| #12 | T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation | Yan Wang, Xinyi Hou, Weiguo Lin +2 | cs.CV | 2026-09-02 |
| #13 | LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails | Vansh Wahi | cs.AI | 2026-09-02 |
| #14 | PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment | Fan Yuxuan, Huang Miaojun, Zhang Haimei +2 | cs.AI | 2026-09-02 |
| #15 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #16 | World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models | Chuhan Zhang, Seiji Ito, Kenta Hoshino +2 | cs.CV | 2026-09-02 |
| #17 | OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations | Yixiong Xiao, Lang An, Hucheng Yang +9 | cs.HC | 2026-09-02 |
| #18 | Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development | Yilong Guo, Hanqi Chen, Zixiao Ye +3 | cs.CV | 2026-09-02 |
| #19 | ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction | Ke Zhang, Yankang Liu, Roya Zandi +1 | cs.AI | 2026-09-02 |
| #20 | HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs | Yihang Sun, Zhihan Zhu, Zhiyuan Jiang +3 | cs.CL | 2026-09-02 |
| #21 | Privacy Washing: Detecting Internal Contradictions in Privacy Policies | Thomas Brackin | cs.CY | 2026-09-02 |
| #22 | Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification | Xuanbing Wen, Boxu Chen, Le Yang +4 | cs.CV | 2026-09-02 |
| #23 | GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric Matching | Lu Han, Xiting Sun, Hao Wang +4 | cs.CV | 2026-09-02 |
| #24 | ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations | Peiying Zhu, Sidi Chang | cs.AI | 2026-09-02 |
| #25 | Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight | Xinyu Fu, Narayan Ramasubbu, Dennis Galletta | cs.HC | 2026-09-02 |
| #26 | Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition | Weiming Li, Catarina Barata, Miguel Constante +1 | cs.CL | 2026-09-01 |
| #27 | InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations | Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3 | cs.CL | 2026-09-01 |
| #28 | IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals | Md. Atabuzzaman, Christian Alexander, Chris Thomas | cs.CV | 2026-09-01 |
| #29 | EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents | Wei Wang, Wenqiao Zhang, Yutong Lin +14 | cs.RO | 2026-09-01 |
| #30 | Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents | Liming Pu, Xiaoxia Li, Yifu Liu +2 | cs.LG | 2026-09-01 |
| #31 | MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence | Walid Saidi | cs.CR | 2026-09-01 |
| #32 | Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges | Rui Yang, Shuang Huang, Junhua Liu +7 | cs.CR | 2026-09-01 |
| #33 | Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set | Michael Zang, Haiyu Wu, Mrinal Sharma +1 | cs.CV | 2026-09-01 |
| #34 | EDRAC: Benchmarking Arabic Dialect Reading Comprehension | Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15 | cs.CL | 2026-09-01 |
| #35 | Causal Evidentiary Governance for High-Risk Machine Learning Systems | Samah Kareem, Barış Çeliktaş | cs.CY | 2026-09-01 |
| #36 | Low-Quality Face Recognition using Center Aligned Representations and Local Margin Constraints | Vedat Can Dilaver, Benjamin S. Riggan | cs.CV | 2026-09-01 |
| #37 | VerNav: Verifier-First Low-Latency Vision-and-Language Navigation | Zhixin Wang, Chengzheyi Yao, Leyuan Liu +2 | cs.RO | 2026-09-01 |
| #38 | Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting | Udo Schlegel, Shubhangi, Gabriel Dax +3 | cs.CV | 2026-09-01 |
| #39 | Agentic programs: an emerging form of scientific software in computational materials science | Yunsung Lim, Haekwan Jeon, Jaesun Kim +2 | cond-mat.mtrl-sci | 2026-09-01 |
| #40 | SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification | Swapnil Bhattacharyya, Mayank Baranwal | cs.AI | 2026-09-01 |
| #41 | Controllable Image Captioning with Prompt-Conditioned Scene Rewards | Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim | cs.CV | 2026-09-01 |
| #42 | Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search | Enrong Pan, Ryan Zhou, Ting Hu | cs.AI | 2026-09-01 |
| #43 | A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss | Suryaansh Jain, Rahasya Barkur, Vishal G +8 | cs.CV | 2026-09-01 |
| #44 | Enoki: Efficient Multi-Level Hallucination Detection | Elisei Rykov, Timur Ionov, Nikolay Ivanov +5 | cs.CL | 2026-09-01 |
| #45 | DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement | Pancheng Niu, Jun Guo, Qiaolin He +2 | cs.LG | 2026-09-01 |
| #46 | CoVer: Conflict-Aware Claim Verification | Shuning Zhang, Dai Shi, Bohao Chu +7 | cs.AI | 2026-09-01 |
| #47 | Validity-Aware Jailbreak Evaluation for Large Language Models | Qilong Wu, Sahil Wadhwa, Pranab Mohanty +2 | cs.AI | 2026-08-31 |
| #48 | MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation | Hangxiao Zhu, Suliu Qin, Zhuoyan Li +3 | cs.CL | 2026-08-31 |
| #49 | TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning | Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab +2 | cs.CL | 2026-08-31 |
| #50 | From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education | Gokhan Dogru, Adrià Martín Mor | cs.CL | 2026-08-31 |