| 1 | WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data | Ji Soo Lee, Xilun Chen, Pierce Chuang +5 | cs.CL | 2026-09-04 |
| 2 | Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models | Wonje Jeung, Sangyeon Yoon, Hyesoo Hong +6 | cs.RO | 2026-09-04 |
| 3 | Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe | Dain Kim, Eungi Cho, Kyumin Kim +2 | cs.AI | 2026-09-04 |
| #4 | Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability | Ankit Goyal, Jaideep Ray | cs.AI | 2026-09-04 |
| #5 | Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models | José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez +2 | cs.AI | 2026-09-04 |
| #6 | How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method | Konstantin Grotov, Valentin Malykh | cs.LG | 2026-09-04 |
| #7 | Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization | Sihan Ge, Yichen Lin, Chenyu Zhou +3 | math.OC | 2026-09-04 |
| #8 | Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC | Saad Memon, Rafal Graczyk, Jan Swakoń +3 | cs.AR | 2026-09-04 |
| #9 | Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory | Peng Cui, Heejin Do, Mrinmaya Sachan | cs.AI | 2026-09-04 |
| #10 | Uncensored Open-weight Models: Redistribution as the Persistence Layer | 10a Labs, :, Juliette Garcia +7 | cs.AI | 2026-09-04 |
| #11 | Substrate-Aware AI Agents: Execution Context as a First-Class Input | Manu Agrawal | cs.AI | 2026-09-04 |
| #12 | Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets | Arunan J | cs.LG | 2026-09-04 |
| #13 | Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers | Yumiao Li, Peixin Liu, Donglin Di +2 | cs.CL | 2026-09-04 |
| #14 | Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG | Shuyu Guo, Shuo Zhang, Zhaochun Ren | cs.CL | 2026-09-04 |
| #15 | A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment | María Eugenia Curi, Germán Capdehourat, Isabel Amigo +4 | cs.CL | 2026-09-04 |
| #16 | NS-ST-GraphRAG: Neuro-Symbolic Spatio-Temporal GraphRAG for Literary Knowledge Processing | Zheng Kui Lin | cs.CL | 2026-09-04 |
| #17 | TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors | Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung +1 | cs.CR | 2026-09-04 |
| #18 | TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents | Zhibo Yang, Chen Zhang, Yuewei Zhang +1 | cs.AI | 2026-09-04 |
| #19 | Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources | Dmitrij Żatuchin | cs.IR | 2026-09-04 |
| #20 | Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications | Yanchen Li, Xiaoming Xue, Kay Chen Tan | cs.AI | 2026-09-04 |
| #21 | BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation | André Ribeiro, Rúben Garrido, Alexander Christiansen +2 | cs.CL | 2026-09-04 |
| #22 | BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference | Janghyeon Kim, Minsoo Kim, Kyuhong Shim +1 | cs.LG | 2026-09-04 |
| #23 | MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision | Zijie Zhu, Weiren Cai, Yizhou Wang +4 | cs.CV | 2026-09-04 |
| #24 | RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents | Aziz Ben Amor, Drish Mali, Mann Acharya +2 | cs.CL | 2026-09-04 |
| #25 | From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments | Linsen Zhu, Mengqing Cai | cs.AI | 2026-09-04 |
| #26 | PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation | Yuchen Sun, Qian Yang, Jun Wang +4 | cs.MM | 2026-09-04 |
| #27 | MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain | Sourav Malakar, Harshit Nigam, Akash Ghosh +4 | cs.CL | 2026-09-04 |
| #28 | Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates | Manuel Röder, Bibin Babu, Frank-Michael Schleif | cs.LG | 2026-09-04 |
| #29 | Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications | Hailiang Zhao, Peng Chen, Xueyan Tang +2 | cs.LG | 2026-09-04 |
| #30 | Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges | Chenqi Li, Minghui Min, Dusit Niyato +1 | cs.AI | 2026-09-04 |
| #31 | Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving | Aditi Patodiya | cs.SE | 2026-09-04 |
| #32 | AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editing | Bo-Han Kung, Futa Waseda, Ching-Chun Chang +2 | cs.CV | 2026-09-04 |
| #33 | Wireless Foundation Models: State-of-the-Art and Open Challenges | Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh +3 | eess.SP | 2026-09-04 |
| #34 | A Differentiable Neural Surrogate for Photon Propagation in Neutrino Telescopes | Felix J. Yu, Berthy T. Feng, Nicholas Kamp +1 | astro-ph.HE | 2026-09-04 |
| #35 | Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle | Happy Bhati | cs.SE | 2026-09-04 |
| #36 | SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds | Esteban Gutiérrez, Lonce Wyse, Frederic Font +1 | cs.SD | 2026-09-04 |
| #37 | SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents | Xin He, Yanlin Wang, Mingwei Liu +3 | cs.SE | 2026-09-03 |
| #38 | From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research | Yakov Pyotr Shkolnikov | cs.AI | 2026-09-03 |
| #39 | A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle | Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer | cs.LG | 2026-09-03 |
| #40 | Efficient Test-Time Adaptation through Human-AI Interaction | Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao +22 | cs.AI | 2026-09-03 |
| #41 | Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving | Ruoyu Yao, Yusen Xie, Qingzhao Liu +5 | cs.CV | 2026-09-03 |
| #42 | Spurious Advantage Hidden in GRPO | Jiamian Wang, Samyadeep Basu, Koustava Goswami +2 | cs.AI | 2026-09-03 |
| #43 | A location-invariant estimator of extremal quantile treatment effects for heavy-tailed distributions | Xin Yu, Shuwei Huang, Jicheng Liu +4 | cs.LG | 2026-09-03 |
| #44 | Unlocking Lossless Speedups in LLMs via Discrete Diffusion | Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham +14 | cs.LG | 2026-09-03 |
| #45 | RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models | Mohammad Mohammadi, Alireza Zarei | cs.LG | 2026-09-03 |
| #46 | Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding | Gaspard Quenard, Takudzwa Togarepi, Damien Pellier +1 | cs.AI | 2026-09-03 |
| #47 | Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities | Biraj Subedi | cs.IR | 2026-09-03 |
| #48 | GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs | Junqing Du, Fernando Ropero, Erkin Turkoz +2 | cs.CV | 2026-09-03 |
| #49 | FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation | Yutian Zhang, Siyuan Ma, Liwen Yang +6 | cs.RO | 2026-09-03 |
| #50 | A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors | Pengxun Li, Litian Zhang, Jianwei Hou +4 | cs.CR | 2026-09-03 |