| 1 | SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment | Qinghua Mao, Wanying Qu, Dadi Guo +8 | cs.AI | 2026-09-02 |
| 2 | Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills | Jianlyu Chen, Yuyang Hu, Hongjin Qian +8 | cs.AI | 2026-09-02 |
| 3 | APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering | Jie Ding, Rui Sun, Xinyuan Zhang +2 | cs.AI | 2026-09-02 |
| #4 | RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution | Zijian Zhao, Sen Li, Xialiang Tong +1 | cs.MA | 2026-09-02 |
| #5 | SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams | Ao Yan, Xin Zhang, Jiawei Du +1 | cs.AI | 2026-09-02 |
| #6 | DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation | Wei Zhang, Hongji Li, Song Sun +4 | cs.LG | 2026-09-02 |
| #7 | OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations | Yixiong Xiao, Lang An, Hucheng Yang +9 | cs.HC | 2026-09-02 |
| #8 | MASkills: Continual Skills Optimization for Multi-Agent LLM Systems | Huaiyuan Yao, Xiaoou Liu, Charles Fleming +2 | cs.AI | 2026-09-02 |
| #9 | Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents | Yanting Yang, Can Jin, Jinman Zhao +6 | cs.LG | 2026-09-02 |
| #10 | How Output Format Confounds Data Quality and Capability in Instruction Tuning | Chengguang Gan, Hanjun Wei, Yunhao Liang +3 | cs.CL | 2026-09-02 |
| #11 | Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis | Arif Hassan Zidan, Yi Pan, Bowen Guo +5 | cs.LG | 2026-09-01 |
| #12 | Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents | Xiaofang Yang, Ziqi Miao, Dianbo Sui +2 | cs.CR | 2026-09-01 |
| #13 | Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement | Haoyang Yan, Min-le Su, Hangfan Zhang +6 | cs.AI | 2026-09-01 |
| #14 | TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution | Ruocan Wei | cs.LG | 2026-09-01 |
| #15 | EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents | Wei Wang, Wenqiao Zhang, Yutong Lin +14 | cs.RO | 2026-09-01 |
| #16 | Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents | Jinqing Zhao, Chengcan Wu | cs.AI | 2026-09-01 |
| #17 | Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents | Liming Pu, Xiaoxia Li, Yifu Liu +2 | cs.LG | 2026-09-01 |
| #18 | REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs | Riyaaz Shaik, Chandru Venkataraman | cs.LG | 2026-09-01 |
| #19 | Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation | Will Badr | cs.SE | 2026-09-01 |
| #20 | HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution | Wen Jiang, Mingmin Chu, Yimeng Tian +6 | cs.LG | 2026-09-01 |
| #21 | DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory | Xincheng Wei, Yifan Ding, Yoshua Li +5 | cs.AI | 2026-09-01 |
| #22 | Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning | Jinhu Qi, Minda Hu, Wentao Zhang +4 | cs.CL | 2026-09-01 |
| #23 | Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding | Chengguang Gan, Yunhao Liang, Hanjun Wei +2 | cs.CL | 2026-09-01 |
| #24 | WiseSpec: Requirements-Driven Agents for Code Generation | Zhao Tian | cs.SE | 2026-09-01 |
| #25 | Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents | Seonghyeon Cho, Chanjun Park | cs.CL | 2026-09-01 |
| #26 | GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting | Mohammad Kian Golkar, Luciano Alves de Oliveira, Mohammad Khanjani | cs.LG | 2026-09-01 |
| #27 | mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers | Timothy Kassis | cs.AI | 2026-08-31 |
| #28 | Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations | Fanyou Wu, Suraj Maharjan, Ainur Yessenalina +3 | cs.AI | 2026-08-31 |
| #29 | Dr. Claw: An AI Scientist Workspace for Vibe Research | Dingjie Song, Hanrong Zhang, Dawei Liu +10 | cs.AI | 2026-08-31 |
| #30 | AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning | Scott Compton, Arjun Nagendran | cs.AI | 2026-08-31 |
| #31 | Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations | Ao Qu, Panagiotis Michelakis, Linyuan Han +8 | cs.AI | 2026-08-31 |
| #32 | Evaluating and Improving LLM Self-Modeling | Siqi Zeng, Andre N. Assis, Rowan Wang | cs.CL | 2026-08-31 |
| #33 | S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation | Xuanle Zhao, Xinyuan Cai, Xiang Cheng +1 | cs.LG | 2026-08-31 |
| #34 | SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment | Chaohui Dang, Zheheng Jiang, James Glasbey +3 | cs.CV | 2026-08-31 |
| #35 | Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions | Rodrigo Almeida, Noelia Otero, Jost Arndt +3 | physics.ao-ph | 2026-08-31 |
| #36 | SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents | Xiaofan Bai, Chao Liu, Hongqiang Lin +5 | cs.AI | 2026-08-31 |
| #37 | PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents | Ziyi Bai, Siqi Li, Tinglei Huang +1 | cs.LG | 2026-08-31 |
| #38 | Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning | Fukang Zhu, Binbin Zhao, Ruixiao Lin +3 | cs.CR | 2026-08-31 |
| #39 | EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents | Doyun Kim, Chanwoo Kim, Sugyeong Eo +2 | cs.AI | 2026-08-31 |
| #40 | From Metaheuristics to Exact Methods: A CP-SAT Approach for Multi-Objective Healthcare Workforce Scheduling | Vipul Patel, Anirudh Deodhar, Dagnachew Birru | cs.AI | 2026-08-31 |
| #41 | Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence | Ziheng Li, Xichen He, Haoyan Chen +10 | cs.AI | 2026-08-31 |
| #42 | Diffusion-Based Refinement for Kilometer-Scale Probabilistic Precipitation Nowcasting | Dohyun Park, Changhoon Song, Tengyuan Chang +2 | cs.LG | 2026-08-31 |
| #43 | Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection | Wujie Xiong, Rabimba Karanjai, Yang Lu +2 | cs.CR | 2026-08-30 |
| #44 | DataFoundry: Evolving Data Preparators via Recursive Self-Improvement | Cehao Yang, Xiaojun Wu, Xueyuan Lin +4 | cs.CL | 2026-08-30 |
| #45 | Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents | Timothy Kassis, Vinayak Agarwal, Yuhuan He +2 | cs.CL | 2026-08-30 |
| #46 | SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs | Yanming Liu, Xinyue Peng, Jiannan Cao +2 | cs.CL | 2026-08-30 |
| #47 | CineForge: Self-Improving Agents for Long-Horizon Video Generation | Junxiang Liu, Lin Wang, Haiyu Shi +10 | cs.CV | 2026-08-30 |
| #48 | Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security | Sanket Badhe, Deep Shah, Priyanka Tiwari +1 | cs.AI | 2026-08-30 |
| #49 | Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation | Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce | cs.CL | 2026-08-30 |
| #50 | AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies | Hongbo Gao, Zeyu Ni, Xin Wen +2 | cs.RO | 2026-08-30 |