| 1 | Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints | Haoyaun Zhu, Jie Zhang | cs.AI | 2026-09-03 |
| 2 | A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms | Davide Paglieri, Logan Cross, Tim Genewein +3 | cs.AI | 2026-09-03 |
| 3 | From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research | Yakov Pyotr Shkolnikov | cs.AI | 2026-09-03 |
| #4 | SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center | Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild | cs.CR | 2026-09-03 |
| #5 | PatchBench: Evaluating AI Agents for Vulnerability Patching | Chihao Shen, Jiacheng Li, Aastha Mahajan +3 | cs.CR | 2026-09-03 |
| #6 | FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models | Yalun Wu, Junfeng Fang, Jiawei Wang +6 | cs.AI | 2026-09-03 |
| #7 | RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models | Mohammad Mohammadi, Alireza Zarei | cs.LG | 2026-09-03 |
| #8 | IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition | Nazim-E-Alam, Tarek Rahman, Md Kishor Morol | cs.CV | 2026-09-03 |
| #9 | Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing | Kubilay Kağan Kömürcü, İlkay Öksüz | cs.CV | 2026-09-03 |
| #10 | Interface-Induced Trajectory Censoring | Wenbo Wang | cs.AI | 2026-09-03 |
| #11 | VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch | WenJie Fan | cs.LG | 2026-09-03 |
| #12 | Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating | Heejin Choi | cs.LG | 2026-09-03 |
| #13 | Xiaomi-TabLDM: A Tabular Foundation Model Technical Report | Xiaomi-TabLDM Team, :, Penghui Wang +10 | cs.AI | 2026-09-03 |
| #14 | Bioinfoysis Technical Report | Qingyang Shao, Xin Zhang, Zhouyang Yuan +24 | cs.AI | 2026-09-03 |
| #15 | Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications | Minwei Zhao, Weiming Zhang, Jiawang Du +4 | cs.CV | 2026-09-03 |
| #16 | DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions | Junjie Pang, Zhenzhen Xie, Haoke Han +3 | cs.AI | 2026-09-03 |
| #17 | RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents | JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong +10 | cs.CV | 2026-09-03 |
| #18 | OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education | Elakkiya Rajasekar | cs.LG | 2026-09-03 |
| #19 | ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation | Javier del Pino, Salvador Rodríguez, Alejandro Garabito +2 | cs.CV | 2026-09-03 |
| #20 | Artificial Intelligence for Energy Optimization in Data Centers | Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah | cs.AI | 2026-09-03 |
| #21 | Observation-Conditioned Latent Energy Priors for Sparse Implicit Neural Shape Completion | Paul Büschl, Ezequiel de la Rosa, Julia Wolleb +3 | cs.CV | 2026-09-03 |
| #22 | MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT | Chenguang Zheng, Le Xue, Yichi Zhang +8 | cs.CV | 2026-09-03 |
| #23 | ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models | Juntao Xu, Shihong Li, Hoi Fan Au +1 | cs.CV | 2026-09-03 |
| #24 | Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements | Arianna Miola, Bruno Spaccavento, Lorenzo Silotto +2 | cs.CL | 2026-09-03 |
| #25 | KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records | Tasmiad Hasan, Arafat Zaman Ratul, Sarker Sadman Saalim +3 | cs.CL | 2026-09-03 |
| #26 | Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation | Haoyu Wang, Songchun Zhang, Haoran Li +3 | cs.CV | 2026-09-03 |
| #27 | NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis | Yinan Liu, Hongtai Xia, Haoran Xu +3 | cs.AI | 2026-09-03 |
| #28 | LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues | Jiayi Li, Zhaomin Wu, Bingsheng He | cs.LG | 2026-09-03 |
| #29 | Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps | Shuning Zhang, Liang Li, Yunheng Wang +3 | cs.RO | 2026-09-03 |
| #30 | AutoGraphForge: Towards Automated Graph Theory Discovery | Ján Pastorek | cs.AI | 2026-09-03 |
| #31 | Mind the Gap: Robustness Risks in PII Detection Systems | Adeel Zafar, Slawomir Nowaczyk | cs.LG | 2026-09-03 |
| #32 | BMCTrack-d: Pig re-identification and tracking via back marks in challenging camera settings | David Brunner, Maciej Oczak, Marie Bordes +3 | cs.CV | 2026-09-03 |
| #33 | Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory | Kazuki Nakayashiki | cs.IR | 2026-09-03 |
| #34 | OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement | Linnan Zhao, Kang Liu, Hao Yu +3 | cs.CV | 2026-09-03 |
| #35 | The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems | Guangjun Liu | cs.MA | 2026-09-03 |
| #36 | Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection | Weijie Liu, Running Zhao, Wenhao Yuan +4 | cs.AI | 2026-09-03 |
| #37 | Time Without Timesteps: Simulating Coupled Dynamical Systems via Self-Consistency | Liyu Zerihun, Mark Shinyoung Lee | cs.LG | 2026-09-03 |
| #38 | Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers | Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan +4 | cs.CL | 2026-09-03 |
| #39 | Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor | Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar +1 | cs.CL | 2026-09-02 |
| #40 | ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers | Ali Hojjat, Janek Haberer, Olaf Landsiedel | cs.CV | 2026-09-02 |
| #41 | MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval | Meriem Yacoubi, Pia Schmidt, Nenad Petrovic +3 | cs.CL | 2026-09-02 |
| #42 | Feasible but Not Safe: Constraint Violations and Report-Channel Attacks in Learned Cell-Free ISAC Association | Mehdi Zafari, Iman Mohammadi, A. Lee Swindlehurst | cs.NI | 2026-09-02 |
| #43 | Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling | Young Seok Jeon, Beatrice Brown-Mulry, Rohan Satya Isaac +5 | cs.CV | 2026-09-02 |
| #44 | IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report] | Lulu Xie, Yancheng Wang, Kanchan Chowdhury +3 | cs.CV | 2026-09-02 |
| #45 | Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains | Yury Korolev | cs.CR | 2026-09-02 |
| #46 | ObserverBench: Testing Mechanistic Estimates for Intervention and Control | Vijay Erramilli | cs.LG | 2026-09-02 |
| #47 | A Common Measure of Communication for Speech Brain-Computer Interfaces | Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones | cs.LG | 2026-09-02 |
| #48 | The Implications of Linguistic Illegibility for LLM Security | James Mickens | cs.LG | 2026-09-02 |
| #49 | UE5M3 FP4 Block Scaling for Stable Language Model Pretraining | Robert Hu, Carlo Luschi, Paul Balanca | cs.LG | 2026-09-02 |
| #50 | Benchmarking RAW and RGB Restoration in Image Signal Processors | Zihao Lu, Radu Timofte, Marcos V. Conde | cs.CV | 2026-09-02 |