| 1 | SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators | Yuncong Yang, Zhengtao Han, Furkan Ozyurt +6 | cs.CV | 2026-09-08 |
| 2 | Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation | Jiacheng Xu, Feng Chen, Xiuneng Xu +1 | cs.LG | 2026-09-08 |
| 3 | Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails | Zhou Yu, Bin Bi, Shiva Kumar Pentyala +8 | cs.AI | 2026-09-08 |
| #4 | DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination | Yankai Fu, Ning Chen, Junkai Zhao +5 | cs.RO | 2026-09-08 |
| #5 | ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR | Tommy Sha, Skylar Zhai, Siqi Zhao | cs.LG | 2026-09-08 |
| #6 | Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents | Wenbo Gao, Zhaomou Song, Zhiyuan Ji +7 | cs.CL | 2026-09-08 |
| #7 | SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation | Linnan Zhao, Xu Liu, Lingling Li +3 | cs.CV | 2026-09-08 |
| #8 | API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces | Jennifer Wang, Joachim Baumann, Daniel E. Ho +1 | cs.AI | 2026-09-08 |
| #9 | Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation | Youngrok Park, Sangmin Bae, Hojung Jung +6 | cs.LG | 2026-09-08 |
| #10 | X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR | Zhiwei Lin, Kaiqi Fu, Rime Wen +5 | cs.SD | 2026-09-08 |
| #11 | Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR | Youngjun Yu, Sanghwan Jang, Hwanjo Yu | cs.LG | 2026-09-08 |
| #12 | SUN: Reaching for Novelty in Reinforcement Learning | Wenyan Yang, Arsenii Mustafin, Dominik Baumann +2 | cs.LG | 2026-09-08 |
| #13 | CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation | Tinghe Ding, Jiahao Li, He Wang | cs.RO | 2026-09-08 |
| #14 | AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction | Feiyu Zhao, Yuetong Li, Chenxi Xiao | cs.RO | 2026-09-08 |
| #15 | Detecting Authorship in Political Texts with Inductive Stylometry | Gennadii Iakovlev, Levente Littvay | cs.CL | 2026-09-08 |
| #16 | SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs | Shengtian Yang, Ziyu Xiong, Yu Li +3 | cs.AI | 2026-09-08 |
| #17 | Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks | Hongbang Yuan, Zhuoran Jin, Yixin Cao | cs.LG | 2026-09-08 |
| #18 | Miles v0.1: Production-Level Post-Training | RadixArk, :, Tom Chen +11 | cs.LG | 2026-09-08 |
| #19 | TV-Regulated OPD: Direction Matters in On-Policy Distillation | Han Xiao, Yifan Niu, Dongyi Liu +2 | cs.LG | 2026-09-08 |
| #20 | Distillation as Probability Transport: Routed On-Policy Distillation | Tianle Xia, Lingxiang Hu, Yiding Sun +6 | cs.LG | 2026-09-08 |
| #21 | Non-Coherent Over-the-Air Federated Learning: Protocol, Convergence, and Device Scheduling | Haifeng Wen, Nicolò Michelusi, Osvaldo Simeone +2 | cs.IT | 2026-09-08 |
| #22 | Human-Centric Image Captioning with Subject-Centered Spatial Understanding | Bozhou Li, Jiahang Zhang, Yue Ding +11 | cs.CV | 2026-09-08 |
| #23 | What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory | Chen Shen | cs.CL | 2026-09-08 |
| #24 | Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems | Yi Ting Shen, Kentaroh Toyoda, Alex Leung | cs.AI | 2026-09-08 |
| #25 | Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTM | Afsara Khan, Austin Rovinski | cs.AR | 2026-09-08 |
| #26 | NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness | NeoHorse Team, Guoliang Cao, Guohao Dai +34 | cs.CL | 2026-09-08 |
| #27 | Snugi-AI-v2 @ eRisk 2026 Task 2: Early Depression Detection via a Learned Stopping Policy with Sustained Confidence Gate | Yuwen Chiu | cs.CL | 2026-09-08 |
| #28 | When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation | Yiwen Qiu, Linjuan Wu, Dingming Li +7 | cs.CL | 2026-09-08 |
| #29 | Observe Before You Alert: Adaptive Driver Alerting with Vision-Language Models | Yuhang Wang, Lingyao Li, Hao Zhou | cs.CV | 2026-09-08 |
| #30 | DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning | Suyog Khanal, Arun Kumar A, Santu Rana | cs.RO | 2026-09-08 |
| #31 | Inference-Time Nash Alignment | Hadi Hosseini, Debmalya Mandal, Duohan Zhang | cs.AI | 2026-09-08 |
| #32 | Automated Design of Inventory Policy with Large Language Models: An Exploratory Study | Fenghua Yang, Preet Baxi, Yi Zhang +4 | cs.AI | 2026-09-08 |
| #33 | Risk-Conditioned Fine-Tuning of Large Language Models | Zixuan Liu, Fangzheng Wu, Brian Summa +1 | cs.LG | 2026-09-08 |
| #34 | Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study | Aayush Patel, Andrzej Ruszczyński | cs.AI | 2026-09-07 |
| #35 | HyCO: A Hybrid Neural Solver for Combinatorial Optimization | Yuheng Li, Di Yang, Haipeng Chen +1 | cs.LG | 2026-09-07 |
| #36 | Streaming Hierarchical Inference with Tabular Foundation Models | Vitor Crista, Afonso Lourenço, Diogo Martinho +1 | cs.LG | 2026-09-07 |
| #37 | Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining | Yongan Yu, Shantam Raj, Jingwei Ni +3 | cs.CL | 2026-09-07 |
| #38 | The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing | Chenguang Wang, Ming Li, Adebayo Braimah +6 | cs.AI | 2026-09-07 |
| #39 | Emergent Charging Coordination in Electric Delivery Fleets | Javier Vales-Alonso, Juan J. Alcaraz | cs.LG | 2026-09-07 |
| #40 | Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions | Zelin Li, Yiyun Su, Matt White +3 | cs.CE | 2026-09-07 |
| #41 | Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best | Kevin Baum, Rūta Binkytė, Felix Jahn | cs.AI | 2026-09-07 |
| #42 | Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding | Yacine El Yamani, Hanna Krasowski, Elena Vanneaux | eess.SY | 2026-09-07 |
| #43 | From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction | Chaemin Jang, Junsik Min, Jaewoo Choi +9 | cs.AI | 2026-09-07 |
| #44 | Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining | Yuhan Wang, Yurou Chen, Hongye Jiang +1 | cs.RO | 2026-09-07 |
| #45 | Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition | Peng Xie | cs.CL | 2026-09-07 |
| #46 | Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy | Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos | cs.RO | 2026-09-07 |
| #47 | Temporal-Causal Inference for Reinforcement Learning via Automata Learning | Jan Corazza, Daniil Kaminskyi, Simon Lutz +4 | cs.LG | 2026-09-07 |
| #48 | Masking Radar Cognition under Adversarial Surveillance: A Distributional Privacy Framework | Sreedevi K, Nandhini K, Anup Aprem +1 | eess.SP | 2026-09-07 |
| #49 | SkillAlign: Aligning Skill Interfaces for LLM-based Agents | Shuo Ren, Xiaomian Kang, Jiajun Zhang | cs.AI | 2026-09-07 |
| #50 | An Auditable Symbolic-RAG-Generative AI Architecture for Goal-Oriented Conversation Orchestration | Ramon Gonzalez, Antonio Diaz | cs.AI | 2026-09-07 |