| 1 | Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning | Yuanteng Chen, Zhilei Liu, Peisong Wang +7 | cs.LG | 2026-09-22 |
| 2 | KwaiMind Technical Report | Junlong Wu, Zijun Li, Yuting Hu +14 | cs.CV | 2026-09-22 |
| 3 | PACT: From Credit Assignment to Critic Alignment | Jiayan Fu, Hang Xu, Yong Zhang +4 | cs.LG | 2026-09-22 |
| #4 | BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception | Zihan Chen, Hengguang Zhou, Yuan Kang +5 | cs.CV | 2026-09-22 |
| #5 | What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation | Kanghui Tian, Siyuan Liu, Tianxiang Jiang +8 | cs.LG | 2026-09-22 |
| #6 | onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction | Lei Yang, Mengyin Liu, Jia Wang +9 | cs.CL | 2026-09-21 |
| #7 | iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs | Ahmed Khaled Khamis, Xiaotong Ji, Hassan Jaber +4 | cs.LG | 2026-09-21 |
| #8 | Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction | Lujia Bao, Qian Chen, Luyao Cheng +14 | eess.AS | 2026-09-21 |
| #9 | ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination | Foundation Model, Li Auto Inc | cs.CV | 2026-09-21 |
| #10 | 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation | Huanxin Sheng, Zhiling Ye, Haonan Wang +3 | cs.LG | 2026-09-21 |
| #11 | MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents | Ruike Cao, Fanyu Zhao, Fugen Yao +6 | cs.LG | 2026-09-21 |
| #12 | Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning | Santi Ram Tiwari, Nihal Naik, Devbrat Pandey +1 | cs.CV | 2026-09-21 |
| #13 | CLOOPD: Closing the Learner Loop in On-Policy Distillation | Keye Zheng, Hanyu Li, Zhan Cheng +1 | cs.LG | 2026-09-21 |
| #14 | ACLArena: Agent Continue Learning in Multi-stage Post-training | Haixin Wang, Xiaoxuan Wang, Junkai Zhang +9 | cs.AI | 2026-09-21 |
| #15 | Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation | Jie Sun, Mao Zheng, Mingyang Song +8 | cs.CL | 2026-09-20 |
| #16 | Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks | Jie Ying, Zhefan Wang, Zihong Chen +11 | cs.LG | 2026-09-20 |
| #17 | One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents | Jie Zhao, Ziyu Jiang, Suhang Zheng +3 | cs.SE | 2026-09-20 |
| #18 | Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World | Kaixiang Yao, Xu Wang, Miao Pan +7 | cs.AI | 2026-09-19 |
| #19 | Calibrating Teacher--Student Discrepancy for On-Policy Distillation | Qiangqiang He, Jin Li, MingCai Chen | cs.AI | 2026-09-18 |
| #20 | On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation | Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub +4 | cs.LG | 2026-09-18 |
| #21 | GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation | Kaichen Zhang, Yuzhong Hong, Junwei Bao +4 | cs.AI | 2026-09-18 |
| #22 | RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning | Yan Yu, Zhengxi Lu, Yizhou Liu +8 | cs.CL | 2026-09-17 |
| #23 | OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher | Damiano Da Col, Maximilian Igl, Peter Karkus +5 | cs.RO | 2026-09-17 |
| #24 | What Does Privileged Information Add to On-Policy Self-Distillation? | XiuYu Zhang, Wei Chow, Junfeng Fang +2 | cs.CL | 2026-09-17 |
| #25 | When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation | Yuxiao Yang, Tianrun Yu, Shangzhe Li +6 | cs.LG | 2026-09-17 |
| #26 | MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution | Loan Bernat, Matthieu Grard, Ariane Herbulot +1 | cs.AI | 2026-09-17 |
| #27 | Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers | Yihao Ai, Weilong Yan | cs.LG | 2026-09-16 |
| #28 | SEA-LION-v4.8: A Technical Report | Ahmed Mohammad Dabeer, Ahn Jeongmi, Anocha Sutaveephamochanon +45 | cs.CL | 2026-09-16 |
| #29 | Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning | Xinxin Song, Siyuan Li, Tingxiong Xiao +1 | cs.AI | 2026-09-16 |
| #30 | Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation | Shiqi Liu, Zeyu He, Letian Tao +9 | cs.LG | 2026-09-15 |
| #31 | Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions | Liu Cao, Xingze Wu, Jingzhi Cui +4 | cs.RO | 2026-09-15 |
| #32 | ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training | Zhihao Zhang, Mingqi Wu, Qiaole Dong +15 | cs.AI | 2026-09-15 |
| #33 | OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation | Chenhao Qiu, Dawei Li, Yechao Zhang +2 | cs.LG | 2026-09-15 |
| #34 | Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning | Xun Xu, Zaixi Zhang | cs.AI | 2026-09-14 |
| #35 | Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation | Daxin Tan, Dehua Tao, Chengxi Deng +2 | eess.AS | 2026-09-14 |
| #36 | Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings | Mingzhou Jiang, Peixi Wu, Hang Cheng +7 | cs.AI | 2026-09-14 |
| #37 | Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models | Shijian Xu, Andrea Miele, Metod Jazbec +3 | cs.LG | 2026-09-14 |
| #38 | Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training | Yuanhao Yue, Qianli Ma, Chengyu Wang +3 | cs.LG | 2026-09-14 |
| #39 | Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition | Yecheng Wu, Song Han, Han Cai | cs.AI | 2026-09-13 |
| #40 | Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation | Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull +4 | cs.CL | 2026-09-13 |
| #41 | Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation | Zhiyu Gui, Kexin Huang, Jia Guo +6 | cs.LG | 2026-09-13 |
| #42 | Data-free On-policy Distillation | Gengsheng Li, Mao Zheng, Mingyang Song +7 | cs.LG | 2026-09-12 |
| #43 | SenseNova-U1.5: Towards Native Unified Visual Intelligence | Haiwen Diao, Jiahao Wang, Chenjing Ding +62 | cs.CV | 2026-09-10 |
| #44 | A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients | Suwan Wu, Yumeng Lin, Pengcheng Yuan +1 | cs.AI | 2026-09-10 |
| #45 | Negative Self-Distillation: Learning to Reason by Avoiding Flaws | Rongcan Pei, Zhepei Wei, Shuyao Xu +3 | cs.CL | 2026-09-10 |
| #46 | KuaiRP Series Role-playing Models Technical Report | Yipeng Wang, Ziwei Zhang, Jiahui Zhang +2 | cs.AI | 2026-09-10 |
| #47 | ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations | Jiawen Wang, Kevin Yao, Khalid Jawed | cs.RO | 2026-09-10 |
| #48 | On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data | Hongyuan Zhang, Xianda Guo, Yanlun Peng +6 | cs.CL | 2026-09-09 |
| #49 | CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts | Naibin Gu, Qingyi Si, Chenxu Yang +5 | cs.LG | 2026-09-09 |
| #50 | Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models | Rui Zhu, Minglong Cao, Chenyu Zhou +2 | math.OC | 2026-09-09 |