| 1 | AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control | Jiabin Qiu, Zixuan Chen, Hongye Cao +3 | cs.AI | 2026-09-24 |
| 2 | PoEM: Predicting RL Outcomes from Existing Policies | Kimia Hamidieh, Giannis Daras, Antonio Torralba | cs.LG | 2026-09-24 |
| 3 | Self-Adaptive VLA for Robust Robot Deployment | Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang +2 | cs.RO | 2026-09-24 |
| #4 | GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS | Saim Rehman, Muhammad Shafique | cs.CV | 2026-09-24 |
| #5 | AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation | Zhiyu Xu, Weilong Yan, Yufei Shi +4 | cs.CV | 2026-09-24 |
| #6 | SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search | Zhongxin Huang, Songyang Li, Renzhe Zhou +5 | cs.IR | 2026-09-24 |
| #7 | Rufus-Air: An Open LLM Post-Training Recipe | Chia-Yuan Chang, Renyuan Cheng, Rui Feng +19 | cs.CL | 2026-09-24 |
| #8 | Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching | Julius Richter, Christoph Boeddeker, Yoshiki Masuyama +4 | eess.AS | 2026-09-24 |
| #9 | Post-Training Leaves Behavioral Shadows on Unrelated Decisions | Ziyang Zhang, Yubin Jing, Yuanhao Zeng +3 | cs.CL | 2026-09-24 |
| #10 | From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents | Xingyu Su, Abhishek Kumar, Qing Ping +5 | cs.AI | 2026-09-24 |
| #11 | AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining | Qingzhuo Wang, Zikun Wei, Zhihua Wei +1 | cs.AI | 2026-09-24 |
| #12 | Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures | Se Un Park, Yutae Kim, Junyoung Park | eess.AS | 2026-09-24 |
| #13 | ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation | Zichong Meng, Chongjian Ge, Chun-Hao P. Huang +2 | cs.CV | 2026-09-24 |
| #14 | Small yet Assistive: Spatially-Aware Post-Training for Low Vision | Rishabh Choudhary, Shreyansh Raj, Umesh Goyal +6 | cs.CV | 2026-09-23 |
| #15 | Temporal Taxation Compounds Under Post-Training Compression of Whisper Models | Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers +1 | cs.CL | 2026-09-23 |
| #16 | Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models | Zehao Liu, Vasant G. Honavar | cs.LG | 2026-09-23 |
| #17 | Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding | Kaiyang Li, Shaobo Han, Yue Tian +1 | cs.SD | 2026-09-23 |
| #18 | Predicting Quantization Price for Selecting PTQ Configurations Before Deployment | Junbin Qiu, Jian Mu, Weitong Zhang +1 | cs.CL | 2026-09-23 |
| #19 | Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs | Haitong Jiang, Chunlin Liu, Yile Wang +1 | cs.CL | 2026-09-23 |
| #20 | InternW0: A Foundational Physical World Model for Efficient Real-World Interactions | Jisong Cai, Yao Mu, Ganlin Yang +23 | cs.RO | 2026-09-23 |
| #21 | When Context Misleads: In-context Learning with Jurisdiction in Large Language Models | Pei-lin Li, Qingle Liu, Junyang Feng +4 | cs.CL | 2026-09-23 |
| #22 | The Capability Manifold and ML Scaling Laws | Syed Ali Raza Zaidi, Maryam Hafeez | cs.LG | 2026-09-23 |
| #23 | Pistis Technical Report | Heyun Chen, Xiaohan Lan, Jiaxi Li +17 | cs.AI | 2026-09-23 |
| #24 | Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction | Jialu Wang, Jianing Deng, Shuqing Luo +6 | cs.LG | 2026-09-23 |
| #25 | CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments | Yuxuan Li, Will Epperson, Wesley Deng +1 | cs.AI | 2026-09-23 |
| #26 | WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps | Abbas Mammadov, Jerry Y. Huang, Justin Lin +5 | cs.LG | 2026-09-22 |
| #27 | OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials Discovery | Thomas Egg, Harry Winston Sullivan, Ellad B. Tadmor +1 | cs.LG | 2026-09-22 |
| #28 | PACT: From Credit Assignment to Critic Alignment | Jiayan Fu, Hang Xu, Yong Zhang +4 | cs.LG | 2026-09-22 |
| #29 | Disaggregated Quantization: Specializing LLM Prefill and Decode | Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi +2 | cs.LG | 2026-09-22 |
| #30 | COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation | Ruike Cao, Fugen Yao, Liang Dong +3 | cs.LG | 2026-09-22 |
| #31 | GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression | Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis | cs.LG | 2026-09-22 |
| #32 | Visual Jev: Accurate and Efficient Decisions from Shared Visual Context | Guanxu Yu, Yuhang Yao | cs.CV | 2026-09-22 |
| #33 | The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance | Rojin Ziaei | cs.AI | 2026-09-22 |
| #34 | Toolcompass: Guiding Tool Trialing, Not Suppressing It | Junlin Fang, Chong Zhang, Do Nguyen-Thanh +3 | cs.AI | 2026-09-22 |
| #35 | Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA | Wenzhi Fang, Nicholas Tzou, Lazar Valkov +1 | cs.AI | 2026-09-22 |
| #36 | VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models | Jiuyi Xu, Qing Jin, Meida Chen +3 | cs.RO | 2026-09-21 |
| #37 | TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks | Bohao Wang, Chenwei Wu, Hang Zou +7 | cs.CL | 2026-09-21 |
| #38 | Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers | Weihang Ding, Junfei Zhan | cs.LG | 2026-09-21 |
| #39 | SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models | Kewei Zhang, Zheng Chen, Haotong Qin +1 | cs.CV | 2026-09-21 |
| #40 | When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs | Yeji Kim, Mi-Young Kim, Randy Goebel | cs.CL | 2026-09-21 |
| #41 | Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement | Qing Yao, Lijian Gao, Qirong Mao | cs.LG | 2026-09-21 |
| #42 | FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding | Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5 | cs.RO | 2026-09-21 |
| #43 | End-to-end Jordanian dialect speech-to-text self-supervised learning framework | Ali A. Safieh, Ibrahim Abu Alhaol, Rawan Ghnemat | cs.CL | 2026-09-21 |
| #44 | SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration | Linhan Luo, Lequan Lin, Dai Shi +3 | cs.LG | 2026-09-21 |
| #45 | MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents | Ruike Cao, Fanyu Zhao, Fugen Yao +6 | cs.LG | 2026-09-21 |
| #46 | Memory vs. Context? Influential Factors of Factual Recall in Language Models | Guilhem Fouilhé, Nicholas Asher, Philippe Muller | cs.CL | 2026-09-21 |
| #47 | P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution | Ningyuan Yang, Yize Li, Pu Zhao +5 | eess.AS | 2026-09-21 |
| #48 | Cost-Accuracy Trade-offs: Neural Operator vs Classical Numerical Solver | Daniel Zhengyu Huang, Andrew M. Stuart | math.NA | 2026-09-21 |
| #49 | ACLArena: Agent Continue Learning in Multi-stage Post-training | Haixin Wang, Xiaoxuan Wang, Junkai Zhang +9 | cs.AI | 2026-09-21 |
| #50 | Measuring the Assistant's Harmlessness Preferences on the User Turn | Jord Nguyen | cs.CL | 2026-09-20 |