| 1 | SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? | Yuqiao Tan, Shizhu He, Jun Zhao +1 | cs.AI | 2026-09-08 |
| 2 | Training-Free Task Vectors for LLM Behavioral Control | Gabriel J. Perin, Lucas Boscaini, André Araujo +1 | cs.LG | 2026-09-08 |
| 3 | Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values | Yuemei Xu, Kexin Xu, Jian Zhou +3 | cs.CL | 2026-09-08 |
| #4 | SignRefine: Adapting Foundational Video Models for Sign Language Generation | Anton Pelykh, Edward Fish, Ozge Mercanoglu Sincan +1 | cs.CV | 2026-09-08 |
| #5 | Compositional Multilingual and Behavioral Attribute Steering | Hyun Gu Kang, Daniil Gurgurov, Tanja Baeumel +2 | cs.CL | 2026-09-08 |
| #6 | Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering | Wenbo Zhang, Zhongxiang Sun, Zhiguang Han +1 | cs.AI | 2026-09-08 |
| #7 | PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion | Peining Zhang, Jinbo Bi | q-bio.BM | 2026-09-08 |
| #8 | LLM Layers Immediately Correct Each Other | Arjun Patrawala, Jiahai Feng, Erik Jones +1 | cs.CL | 2026-09-07 |
| #9 | From Echo Chambers to Epistemic Monoculture: Large Language Models Present Temporally Contingent Partisan Alignments as Knowledge | Wend K. Tam | cs.CL | 2026-09-07 |
| #10 | The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs | Ziyue Feng, Hongbo Fang, James A. Evans | cs.CL | 2026-09-07 |
| #11 | TrojanWorld: Backdooring World-Model Agents via Imagination Steering | Wenkai Huang, Siyuan Liang, Gaolei Li +4 | cs.LG | 2026-09-07 |
| #12 | Disentangling Steering Vectors | Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi +1 | cs.LG | 2026-09-07 |
| #13 | Steering Interference Reflects the Model's Defaults, Not the Behavior Directions | Srikanth Malla, Chiho Choi, Joon Hee Choi | cs.LG | 2026-09-07 |
| #14 | AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering | Shuhe Wang, Lachlan Cowley, Eduard Hovy +1 | cs.CL | 2026-09-06 |
| #15 | LATS: Levy Adaptive Tree Sampling for Feedback-Driven Diverse Target Discovery | Binglin Ji, Anindya Sarkar, Hengchang Lu +3 | cs.LG | 2026-09-06 |
| #16 | Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs | Saurav Bhandari, Benjamin Wade | cs.LG | 2026-09-06 |
| #17 | Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space | Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala +1 | cs.CL | 2026-09-06 |
| #18 | AutoKD: Autonomous Knowledge Discovery | Qinwen Ge, Bo Ni, Haowei Fu +3 | cs.AI | 2026-09-06 |
| #19 | Can Activation Steering Capture Multidimensional Authorship Style? | Hieu Tran, Calvin Bao, Marine Carpuat | cs.CL | 2026-09-04 |
| #20 | Locating and Steering Refusal Beyond Attention | Preethi Carmel Bosco, Gopalakrishnan Srinivasan | cs.LG | 2026-09-04 |
| #21 | A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle | Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer | cs.LG | 2026-09-03 |
| #22 | Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness | Hoang Cuong Nguyen, Mark Dras, Usman Naseem | cs.CL | 2026-09-03 |
| #23 | Large Language Models in Resolving Contextual Knowledge Conflicts | Xinye Yang, Zhenyang Liu, Ruisi Li +1 | cs.CL | 2026-09-02 |
| #24 | ObserverBench: Testing Mechanistic Estimates for Intervention and Control | Vijay Erramilli | cs.LG | 2026-09-02 |
| #25 | Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance | Sai Niranjan Ramachandran, Suvrit Sra | cs.LG | 2026-09-02 |
| #26 | Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training | Jian Gao, Xiao Zhang, Xun Zhu +2 | cs.AI | 2026-09-02 |
| #27 | IDEEA: training-free Input-Dependent stEEring via Activation cluster matching | Zheng Wang, Muchen Li, Renjie Liao +1 | cs.CL | 2026-09-02 |
| #28 | GAPS: Dimension-Level Gates for Conditional Activation Steering | Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad +1 | cs.CL | 2026-09-01 |
| #29 | What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models | Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3 | cs.CV | 2026-09-01 |
| #30 | Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents | Xiaofang Yang, Ziqi Miao, Dianbo Sui +2 | cs.CR | 2026-09-01 |
| #31 | StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions | Chao Gao, Haijiang Liu, Qiyuan Li +3 | cs.CL | 2026-09-01 |
| #32 | Lagged Coupling: Internal Representations Become Readable Before They Become Causal | Xining Xun | cs.CL | 2026-09-01 |
| #33 | SFAD: Speculative Factuality-Aware Decoding | Guanqiao Chen, Di Wang, Lijie Hu | cs.CL | 2026-09-01 |
| #34 | How Do Language Models Choose Between Context and Memory? | Benjamin Shih, John Winnicki, Arianna Cao | cs.LG | 2026-09-01 |
| #35 | Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning | Zixuan Wang, Yixin Hu, Wen Li +4 | cs.CV | 2026-09-01 |
| #36 | Investigating Assistant Bias in LLM User Simulators Using a Role Vector | Daeheon Jeong, Yoonjoo Lee, Eugene Choi +2 | cs.CL | 2026-09-01 |
| #37 | Topological Steering | Benoît Guérand, Tan Minh Nguyen | cs.LG | 2026-09-01 |
| #38 | Potential-Guided Particle Steering for Negation-Constrained Dexterous Grasping | Geonho Kim, SooGon Kim, Jongmin Lee | cs.RO | 2026-09-01 |
| #39 | Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict | Jungyeon Lee, Yejin Yoon, Taeuk Kim | cs.CL | 2026-09-01 |
| #40 | Latent Mechanisms of Language Control in Multilingual Language Models | Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly | cs.CL | 2026-08-31 |
| #41 | Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning | Yu Yuan, Yaoyou Fan, Lili Zhao +5 | cs.CL | 2026-08-31 |
| #42 | Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization | Shiyun Wa, Yifei Wang, Anna G. Green +2 | cs.LG | 2026-08-31 |
| #43 | Asymmetries in Spontaneous and Instructed Deception | Josiah Luikham | cs.AI | 2026-08-31 |
| #44 | Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation | Tianyu Gao, Zhikai Su, Jiashu Li +5 | cs.LG | 2026-08-31 |
| #45 | Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering | Kirill Bunin, Dmitry Bylinkin, Vladimir Aletov +3 | cs.LG | 2026-08-31 |
| #46 | Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer | Minju Song, Hyeon Hwang, Junhyun Lee +1 | cs.CL | 2026-08-31 |
| #47 | Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents | Yunseok Lee, Yunji Kim, Woojin Lee | cs.AI | 2026-08-31 |
| #48 | Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering | Jin Gan, Xin Li, Jun Luo | cs.CL | 2026-08-31 |
| #49 | ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives | Hoejoon Kwon, Byeonggeuk Lim, Kahyeon Kim +1 | cs.CL | 2026-08-31 |
| #50 | Interpreting and Steering for Safe and Correct Code Generation | Hao Yan, Ziyu Yao | cs.AI | 2026-08-30 |