| 1 | HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning | Xinglong Luo, Yuding Zhang, Yuheng Kuang +5 | cs.LG | 2026-09-25 |
| 2 | LandscapeSHAP: Which Persistent Homology Class Gets the Credit? | Nikola Milićević | math.AT | 2026-09-25 |
| 3 | SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization | Xinyi Ke, Kai Li, Junliang Xing +2 | cs.AI | 2026-09-25 |
| #4 | ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning | Zhenlong Dai, Xujie Song, Zitong Wang +7 | cs.CL | 2026-09-25 |
| #5 | Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More | Deng Pan, Joe Germino, Yihong Ma +4 | cs.AI | 2026-09-25 |
| #6 | Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows | Jinfeng Xu, Zheyu Chen, Ziyue Peng +6 | cs.AI | 2026-09-25 |
| #7 | Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation | Shengjun Zhang, Tingyi Liu, Dong Xie +3 | cs.LG | 2026-09-25 |
| #8 | From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation | Mengdan Zhu, Yufan Zhao, Yao Zhao +5 | cs.IR | 2026-09-24 |
| #9 | AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation | Zhiyu Xu, Weilong Yan, Yufei Shi +4 | cs.CV | 2026-09-24 |
| #10 | IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis | Xingyu Wu, Yuchen Yan, Zhengxi Lu +9 | cs.CL | 2026-09-24 |
| #11 | Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams | Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh +2 | cs.CL | 2026-09-24 |
| #12 | When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration | Hengzhi Ye | cs.AI | 2026-09-24 |
| #13 | CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment | Ruochen Jiao, Besnik Fetahu, Zhenyu Shi +1 | cs.AI | 2026-09-24 |
| #14 | SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL | Yan Zhan, Shaobo Liu, Qiunan Liu +7 | cs.AI | 2026-09-24 |
| #15 | When Does Action Credit Need Updating? | Hongye Yang, Boxiao Huang | cs.AI | 2026-09-24 |
| #16 | Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning | Xincheng Yao, Haobo Fu, Weiming Liu +1 | cs.AI | 2026-09-24 |
| #17 | Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency | Geng Chen, Ruotong Pan, Zhirui Yang +9 | cs.AI | 2026-09-23 |
| #18 | When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment | Jie Zhang, Jingxiao Yang, Zhehao Huang +2 | cs.LG | 2026-09-23 |
| #19 | Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance | Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Diaz, Pablo Delgado Romero | cs.MA | 2026-09-23 |
| #20 | PCQC: Privileged Counterfactual Question Credit for Multi-Turn Medical Dialogue | Chenxuan Li, Jiayi Wan, Xinrong Chen +3 | cs.LG | 2026-09-23 |
| #21 | FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints | Sultan Amed, Tanmay Sen, Sayantan Banerjee | stat.ML | 2026-09-23 |
| #22 | ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning | Ming Ma, Yi Zhu, Yiran Zhong +7 | cs.LG | 2026-09-23 |
| #23 | Pistis Technical Report | Heyun Chen, Xiaohan Lan, Jiaxi Li +17 | cs.AI | 2026-09-23 |
| #24 | MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression | Ke Wan, Yifan Wang, Liheng Lai +1 | cs.CL | 2026-09-23 |
| #25 | Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents | Yefan Zhou, Yang Li, Zeyu Leo Liu +2 | cs.AI | 2026-09-23 |
| #26 | Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams | Shujian Gao, Jiamei Yan, Yuchen Yang +6 | cs.LG | 2026-09-23 |
| #27 | XLOG: A CUDA-Native Engine for Neurosymbolic Integration | Levi Dubrovin, Nikita Pospelov, Kirill Sabitov | cs.AI | 2026-09-23 |
| #28 | Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning | Yuqing Zhou, Hong Wang, Manqing Mao +8 | cs.CL | 2026-09-22 |
| #29 | Reinforcement Learning with Decomposed Subtasks | Mattie Terzolo, Mikolaj Sacha, Ayan Sinha +1 | cs.AI | 2026-09-22 |
| #30 | When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations | Nithin Raghava Ramachandra Narla | cs.LG | 2026-09-22 |
| #31 | Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen | Om Nepal, Sushant Aryal, Oluseyi Olukola +1 | cs.SE | 2026-09-22 |
| #32 | PACT: From Credit Assignment to Critic Alignment | Jiayan Fu, Hang Xu, Yong Zhang +4 | cs.LG | 2026-09-22 |
| #33 | Protocol before progress: leakage-aware evaluation of AIS trajectory prediction | Zobeir Raisi, Vali Mohammad Nazarzehi Had | cs.LG | 2026-09-22 |
| #34 | DefaultGNN: A Dual-Perspective GNN Framework for Predicting Corporate Default from Buyer-Seller Transaction Networks | Junghoon Kim, Hyunsung Kim, Seungyoon Choi +4 | cs.LG | 2026-09-22 |
| #35 | TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks | Bohao Wang, Chenwei Wu, Hang Zou +7 | cs.CL | 2026-09-21 |
| #36 | The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora | John J. O'Hare | cs.CL | 2026-09-21 |
| #37 | Credit Access is Associated with Improved Food Security in the Horn of Africa | Jordi Cerdà-Bautista, Vasileios Sitokonstantinou, José Manuel Veiga López-Peña +3 | cs.LG | 2026-09-21 |
| #38 | Information-Time Proximal Policy Optimization | Yongcheng Zeng, Xinyu Cui, Yan Song +9 | cs.LG | 2026-09-21 |
| #39 | When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain | Yunxiang Li, Xixin Wu, Helen Meng | cs.AI | 2026-09-21 |
| #40 | MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents | Ruike Cao, Fanyu Zhao, Fugen Yao +6 | cs.LG | 2026-09-21 |
| #41 | CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment | Xuchun Hu | cs.AI | 2026-09-21 |
| #42 | Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits | Ali Hedayati Pirouzan | cs.AR | 2026-09-21 |
| #43 | FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model | Jingxuan Xu, Gang Wu, Yanan Wu +13 | cs.CL | 2026-09-20 |
| #44 | Which Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation | Haoyue Liu, Ye Chen, Zhichao Wang +3 | cs.SD | 2026-09-20 |
| #45 | When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used | Kunyu Peng, Junming Liu, Ruiqi He +3 | cs.AI | 2026-09-19 |
| #46 | ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control | Huaitao Zhao, Tianlong Zhou, Weijie Wang +2 | cs.AI | 2026-09-19 |
| #47 | EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise | Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid | cs.AI | 2026-09-18 |
| #48 | ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL | Qiang Zhang, Ruixue Ding, Fanrui Zhang +9 | cs.CL | 2026-09-18 |
| #49 | Ownership in AI-Assisted Everyday Tasks | Megan Wei, Melanie Subbiah, Audrey Lee +4 | cs.AI | 2026-09-17 |
| #50 | greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI | Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh +1 | cs.DL | 2026-09-17 |