| 1 | EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction | Yuling Shi, Zhensu Sun, Junsen Dong +3 | cs.CL | 2026-09-02 |
| 2 | DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models | Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi +7 | cs.CL | 2026-09-02 |
| 3 | From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs | Urja Pawar, Rajitha Ramanayake, Owen O'Neill +4 | cs.CL | 2026-09-02 |
| #4 | When Persona Attributes Improve Population Alignment in Large Language Models | Leon Fröhling, Jens Rupprecht, Markus Strohmaier +1 | cs.CL | 2026-09-02 |
| #5 | PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation | Smitha Muthya Sudheendra, Jaideep Srivastava | cs.CL | 2026-09-02 |
| #6 | Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions | Jiayi Bi, Yanjie Gao, Yuanmin Xie +4 | cs.AI | 2026-09-02 |
| #7 | Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions | Maitreyee Tewari, Michele Persiani | cs.HC | 2026-09-02 |
| #8 | LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails | Vansh Wahi | cs.AI | 2026-09-02 |
| #9 | PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks | Yuyao Zheng, Haipeng Sun, Junwei Bao +4 | cs.AI | 2026-09-02 |
| #10 | Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents | Jalal Mahmud | cs.AI | 2026-09-02 |
| #11 | Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap | Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2 | cs.CV | 2026-09-02 |
| #12 | Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision | Sitong Pan, Yipeng Shen, Yilin Lu +3 | cs.AI | 2026-09-02 |
| #13 | The Dynamics of Continuous Mixture Collapse in Language Models | Ali Backour | cs.LG | 2026-09-02 |
| #14 | Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents | Yanting Yang, Can Jin, Jinman Zhao +6 | cs.LG | 2026-09-02 |
| #15 | Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight | Xinyu Fu, Narayan Ramasubbu, Dennis Galletta | cs.HC | 2026-09-02 |
| #16 | The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents | Jundong Hu, Shekar Ramachandran | cs.AI | 2026-09-01 |
| #17 | Agent Memory Is a Surface for Endogenous Authorization Laundering | Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol | cs.CR | 2026-09-01 |
| #18 | VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages | Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda +4 | cs.CL | 2026-09-01 |
| #19 | Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation | Ehsan Faghih, Fatemeh Ashrafi, Marguerite Moore +1 | cs.CV | 2026-09-01 |
| #20 | MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models | Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5 | cs.CL | 2026-09-01 |
| #21 | Evidential Deep Learning for Multi-Modal Anti-UAV Detection | Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag | cs.CV | 2026-09-01 |
| #22 | SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition | Biraj Subedi | cs.CL | 2026-09-01 |
| #23 | Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation | Kefeng Duan, Dewu Zheng, Yanlin Wang +8 | cs.SE | 2026-09-01 |
| #24 | Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation | Haoyuan Deng, Haichao Liu, Wenkai Guo +6 | cs.RO | 2026-09-01 |
| #25 | From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix | Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov +11 | cs.CL | 2026-09-01 |
| #26 | Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories | Nabira Rashid, Manolis Kellis | cs.LG | 2026-09-01 |
| #27 | When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation | Peiying Zhu, Sidi Chang | cs.AI | 2026-09-01 |
| #28 | FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making | Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza | cs.CV | 2026-09-01 |
| #29 | Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers | Egor Pakhomov, Erik Nijkamp | cs.AI | 2026-09-01 |
| #30 | When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning | Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2 | cs.CR | 2026-09-01 |
| #31 | Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds | Clinton Enwerem, John S. Baras, Calin Belta | cs.RO | 2026-09-01 |
| #32 | EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems | Jun Hou, Priya Pitre, Yi Fang +1 | cs.AI | 2026-09-01 |
| #33 | Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR | Esther Xin | cs.CL | 2026-09-01 |
| #34 | CHARM: Character Hallucination for Multicultural Role Play Benchmark | Sunkyung Han, Nahyeon Park, Gaeun Seo +2 | cs.CL | 2026-09-01 |
| #35 | ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence | Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao +1 | cs.CV | 2026-09-01 |
| #36 | Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents | Ruochen Zhou, Zhengyu Chen, Luan Zhang +3 | cs.CL | 2026-09-01 |
| #37 | HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives | Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera +1 | cs.CV | 2026-09-01 |
| #38 | What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal | Radin Shayanfar, Keheliya Gallaba, Ahmed E. Hassan | cs.SE | 2026-09-01 |
| #39 | MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation | Zhijian Qiao, Xinjiang Wang, Jiajie Chen +5 | cs.CV | 2026-09-01 |
| #40 | One Prompt Is Enough: Watermark Laundering Through Foundation Image Models | Jidong Yang, Qi Li, Wei Zong +5 | cs.CV | 2026-09-01 |
| #41 | Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents | Liming Pu, Xiaoxia Li, Yifu Liu +2 | cs.LG | 2026-09-01 |
| #42 | MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence | Walid Saidi | cs.CR | 2026-09-01 |
| #43 | Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening | Zhilong Song, Lixue Cheng | cond-mat.mtrl-sci | 2026-09-01 |
| #44 | Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation | Will Badr | cs.SE | 2026-09-01 |
| #45 | When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP | Shota Sato, Hajime Kiyama, Tosho Hirasawa +1 | cs.CL | 2026-09-01 |
| #46 | Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models | Jiayu Ding, Zhuodong Liu, Lei Zhang +6 | cs.CV | 2026-09-01 |
| #47 | Disclosure-Gated User Simulation for Companion-Agent Evaluation | Yao Liu, Yu He | cs.CL | 2026-09-01 |
| #48 | Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA | Hai-Dang Nguyen, Huy-Hieu Pham | cs.CV | 2026-09-01 |
| #49 | Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling | Kangjia Zhao, Jiajun Li, Haozhan Shen +8 | cs.CL | 2026-09-01 |
| #50 | A Dataset for Modeling Iterative Problem-Solving | Fagun Patel, Sang T. Truong, Duc Q. Nguyen +4 | cs.CL | 2026-09-01 |