| 1 | Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework | Cagri Temel | cs.RO | 2026-09-02 |
| 2 | SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment | Qinghua Mao, Wanying Qu, Dadi Guo +8 | cs.AI | 2026-09-02 |
| 3 | CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation | Varun Gadey, Ziad Marey, Alexandra Dmitrienko | cs.CR | 2026-09-02 |
| #4 | From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution | Yuzhang Luo, Chenpeng Wang, Jianhui Chen +1 | cs.CL | 2026-09-02 |
| #5 | SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective | James Di Novo, Hany Ragab, Sylvain P. Leblanc | cs.CR | 2026-09-02 |
| #6 | Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking | Siyu Chen, Haoran Wang, Xiaojian Li +3 | cs.CL | 2026-09-02 |
| #7 | Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization | Bing Zheng, Zongyao Zhao, Wenming Yang | cs.IR | 2026-09-02 |
| #8 | Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds | Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6 | cs.AI | 2026-09-02 |
| #9 | SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment | Qingyu Meng, Yiwei Zha, Jiahuan Pei +3 | cs.LG | 2026-09-02 |
| #10 | SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology | Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko +2 | cs.AI | 2026-09-02 |
| #11 | If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection | Yinghao Sun, Shuguang Li, Jinliang Shao +1 | cs.CV | 2026-09-02 |
| #12 | CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging | Mingjie Zheng, Zihao Chen, Wenqing Chen +4 | cs.AI | 2026-09-02 |
| #13 | CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation | Shucheng Zhang, Yuang Zhang, Bingzhang Wang +3 | cs.RO | 2026-09-02 |
| #14 | ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models | Da Cheng Gu, Yifei Dong, Xinghao Yang +2 | cs.AI | 2026-09-02 |
| #15 | FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs | Zhengyi Jin, Ru Zhang, Xiao Chen +5 | cs.AI | 2026-09-02 |
| #16 | Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage | Weifeng Jiang, Ruirui Chen, Qianren Mao +3 | cs.CL | 2026-09-02 |
| #17 | Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models | Tianqi Xiao, Shiyao Cui, Minghao Zhang +2 | cs.MM | 2026-09-02 |
| #18 | Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets | Kunal Jadhav, Siddhesh More | cs.CL | 2026-09-01 |
| #19 | The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents | Jundong Hu, Shekar Ramachandran | cs.AI | 2026-09-01 |
| #20 | Agent Memory Is a Surface for Endogenous Authorization Laundering | Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol | cs.CR | 2026-09-01 |
| #21 | SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue | Stephanie Fong, Yiwen Jiang, Zimu Wang +12 | cs.CL | 2026-09-01 |
| #22 | Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study | Arpan Kumar Mahapatra | cs.CR | 2026-09-01 |
| #23 | GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions | Elias Stengel-Eskin, Newton Sander, Carlos Bonetti +4 | cs.CL | 2026-09-01 |
| #24 | Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents | Xiaofang Yang, Ziqi Miao, Dianbo Sui +2 | cs.CR | 2026-09-01 |
| #25 | RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching | Charles Corbière, Léo Machado, Aubin Charley +3 | cs.CV | 2026-09-01 |
| #26 | When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning | Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2 | cs.CR | 2026-09-01 |
| #27 | Provably Safe Sim-to-Real Transfer | Tingting Ni, Maryam Kamgarpour | cs.LG | 2026-09-01 |
| #28 | The Constitutional Coverage Trilemma in AI Governance | Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly +1 | cs.LG | 2026-09-01 |
| #29 | Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges | Rui Yang, Shuang Huang, Junhua Liu +7 | cs.CR | 2026-09-01 |
| #30 | Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate | Kaiyan Wen, Shijie Zhang, Lu Yu +1 | cs.AI | 2026-09-01 |
| #31 | HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation | Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin | cs.CR | 2026-09-01 |
| #32 | Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees | Molly Wang | cs.AI | 2026-09-01 |
| #33 | In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access? | Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi +2 | cs.AI | 2026-09-01 |
| #34 | A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals | Yuri Son, Seunghee Kim, Hyuhng Joon Kim +1 | cs.CL | 2026-09-01 |
| #35 | Patterning in Practice: Debiasing Reward Models with Susceptibilities | George Wang, Elizabeth Donoway, Daniel Murfet | cs.LG | 2026-09-01 |
| #36 | Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs | Jainil Dharmil Shah | cs.AI | 2026-09-01 |
| #37 | Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time | Zeen Zhu, Zhuo Li, Weiyang Guo +4 | cs.CL | 2026-09-01 |
| #38 | Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing | Alexandre Clin Deffarges, Nataliya Kosmyna, Pattie Maes | cs.AI | 2026-09-01 |
| #39 | Validity-Aware Jailbreak Evaluation for Large Language Models | Qilong Wu, Sahil Wadhwa, Pranab Mohanty +2 | cs.AI | 2026-08-31 |
| #40 | Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models | Guoli Wang, Haonan Shi, Tu Ouyang +1 | cs.CL | 2026-08-31 |
| #41 | EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities | Feitong Qiao, Liren Peng, Shiming Ren +7 | cs.CL | 2026-08-31 |
| #42 | Capability-Gated Language Models: Security Composes, Utility Does Not | Patrikas Vanagas, Augustas Mačijauskas, Laurynas Lopata | cs.CR | 2026-08-31 |
| #43 | Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You | Salim Khazem, Ibrahim Mohamed Serouis | cs.LG | 2026-08-31 |
| #44 | The Answer Is Not the Argument | Will Yeadon, Sergio Juárez, Paul Mackay +5 | cs.AI | 2026-08-31 |
| #45 | Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation | Ruotong Wang, Zihao Zhu, Siwei Lyu +2 | cs.CV | 2026-08-31 |
| #46 | Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving | Zhengxu Tang, Xiaozhou Zhang, Guofeng Cui +10 | cs.CV | 2026-08-31 |
| #47 | BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing | Adrians Skapars, Edoardo Manino | cs.AI | 2026-08-31 |
| #48 | Generative artificial intelligence for reliable mechanistic reasoning for corrosion | Bharath M N, R K Singh Raman, Alankar Alankar | cs.LG | 2026-08-31 |
| #49 | TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning | Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick +3 | cs.CL | 2026-08-31 |
| #50 | Safety Screening for Voltage Control in Active Distribution Grids via Distributionally Robust Conformal Screening | Sarra Bouchkati, Petros Ellinas, Adriana Geisler +4 | eess.SY | 2026-08-31 |