| 1 | Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization | Giovanni Dispoto, Marcello Restelli, Carmine Ventre | q-fin.PM | 2026-09-02 |
| 2 | DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models | Yotam Eshel, Guy Hadad, Guy Feigenblat +3 | cs.LG | 2026-09-02 |
| 3 | Fair Stable Matching: A Nash Social Welfare Approach | Parth Desai, Rasheed M, Ganesh Ghalme +1 | cs.GT | 2026-09-02 |
| #4 | RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification | Jierui Li, Zhiyuan Qi, Hao Zhu +7 | cs.CV | 2026-09-02 |
| #5 | CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging | Mingjie Zheng, Zihao Chen, Wenqing Chen +4 | cs.AI | 2026-09-02 |
| #6 | CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents | S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol | cs.LG | 2026-09-02 |
| #7 | Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality | Yifan Zhu, Sammie Katt, Samuel Kaski | cs.AI | 2026-09-02 |
| #8 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #9 | GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation | Qianqian Wang, Yunshan Li, Jiawen Zeng +2 | cs.IR | 2026-09-02 |
| #10 | Thinking effort aligns between humans and reasoning models in abductive reasoning | Henry Arthur | cs.CL | 2026-09-01 |
| #11 | The Rise of Verbal Reinforcement Learning | Kshitij Tayal, Arun Sharma, Genta Indra Winata +2 | cs.CL | 2026-09-01 |
| #12 | Mechanism Design for Alignment and Control | Dirk Bergemann, Andrew Koh, Stephen Morris | econ.TH | 2026-09-01 |
| #13 | Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs | Jingtan Wang, Arun Verma, Xiaoqiang Lin +4 | cs.CL | 2026-09-01 |
| #14 | Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation | Thibaut Thonet, Jos Rozen, Laurent Besacier | cs.CL | 2026-09-01 |
| #15 | CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs | Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3 | cs.LG | 2026-09-01 |
| #16 | Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation | Zhixuan Liu, Zhichen Dong, Yuyu Fan +2 | cs.LG | 2026-09-01 |
| #17 | Does This Moment Justify the Recommendation? Counterfactual Behavior-Grounded Evidence Retrieval for Personalized Video Recommendation | Xin Liu | cs.CV | 2026-09-01 |
| #18 | VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences | Yiwen Jiang, Yang Deng, Stephanie Fong +9 | cs.AI | 2026-09-01 |
| #19 | VerNav: Verifier-First Low-Latency Vision-and-Language Navigation | Zhixin Wang, Chengzheyi Yao, Leyuan Liu +2 | cs.RO | 2026-09-01 |
| #20 | RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation | Zhongru Chen, Yuan Wu, Yi Chang | cs.AI | 2026-09-01 |
| #21 | A multicenter benchmark and clinically structured metric for coronary CTA report generation | Zhiyu Ye, Yue Sun, Limiao Zou +7 | cs.CV | 2026-09-01 |
| #22 | When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting | Ariel Smogorghevski, Nir Rosenfeld, Yaniv Romano | cs.LG | 2026-09-01 |
| #23 | Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning | Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee +2 | cs.AI | 2026-09-01 |
| #24 | CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction | Kewei Li, Rongying Zhang, Peiyu Yang +4 | cs.LG | 2026-09-01 |
| #25 | Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources | Ivan Decostanzi, Michele Ronco, Sergio Consoli +8 | cs.AI | 2026-09-01 |
| #26 | SFAD: Speculative Factuality-Aware Decoding | Guanqiao Chen, Di Wang, Lijie Hu | cs.CL | 2026-09-01 |
| #27 | Instella-MoE Technical Report | Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10 | cs.CL | 2026-09-01 |
| #28 | Patterning in Practice: Debiasing Reward Models with Susceptibilities | George Wang, Elizabeth Donoway, Daniel Murfet | cs.LG | 2026-09-01 |
| #29 | Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time | Zeen Zhu, Zhuo Li, Weiyang Guo +4 | cs.CL | 2026-09-01 |
| #30 | Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search | Jincheng Zhang, Chen Huang, Wenqiang Lei +2 | cs.IR | 2026-09-01 |
| #31 | Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict | Jungyeon Lee, Yejin Yoon, Taeuk Kim | cs.CL | 2026-09-01 |
| #32 | Emotional Labor Strategy Preferences in LLM Personas | Mohammad Saim, Tianyu Jiang | cs.CL | 2026-08-31 |
| #33 | The Assistant's Ideal Self | Mert Yazan | cs.AI | 2026-08-31 |
| #34 | Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching | Kartik Ravisankar, Hojat Abdolanezhad, Daniel Capo +3 | cs.AI | 2026-08-31 |
| #35 | Hypotheses-Guided Self Distillation for Continual Personalization | EunJeong Hwang, Kushan Mitra, Dan Zhang +2 | cs.AI | 2026-08-31 |
| #36 | Authority Bias in Conversational Search Engines for Academic Paper Recommendation | Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki +3 | cs.AI | 2026-08-31 |
| #37 | Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization | Camila Blank, Zhuofan Ying, Christopher Potts +2 | cs.LG | 2026-08-31 |
| #38 | Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization | Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang +3 | cs.AI | 2026-08-31 |
| #39 | TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning | Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick +3 | cs.CL | 2026-08-31 |
| #40 | Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation | Alessio Galatolo, Meriem Beloucif | cs.CL | 2026-08-31 |
| #41 | PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation | MinKeon Kim, Namjun Lee, Jaekwang Kim | cs.CL | 2026-08-31 |
| #42 | Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments | Ivan Porupski, Nikola Ljubešić | cs.CL | 2026-08-31 |
| #43 | Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval | Shaowei Wei, Chong Huang, Songtao Fang +3 | cs.IR | 2026-08-31 |
| #44 | Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment | Yifan Zhu, Kyeongmin Rim, James Pustejovsky | cs.CL | 2026-08-31 |
| #45 | PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization | Boryeong Cho, Sumyeong Ahn, Se-Young Yun | cs.LG | 2026-08-31 |
| #46 | Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval | Gaoming Zhang, Angqing Jiang, Jianchun Song +4 | cs.IR | 2026-08-31 |
| #47 | From Metaheuristics to Exact Methods: A CP-SAT Approach for Multi-Objective Healthcare Workforce Scheduling | Vipul Patel, Anirudh Deodhar, Dagnachew Birru | cs.AI | 2026-08-31 |
| #48 | Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation | Jinyoung Kim, Muhammad Khalifa, Lajanugen Logeswaran +4 | cs.CL | 2026-08-31 |
| #49 | Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS | Yan Zhou, Yun Hong, Yang Feng | cs.CL | 2026-08-31 |
| #50 | Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs | Shunjie Wen, Jaeyeon Lee, Dong-Wan Choi | cs.CV | 2026-08-31 |