| 1 | User Feedback Provides a Unique Signal that LLMs Can not Detect | Shachar Don-Yehiya, Leshem Choshen, Omri Abend | cs.CL | 2026-09-02 |
| 2 | Post-Training Language Models for Gold-Medal Performance in Coding Competitions | Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi +2 | cs.LG | 2026-09-02 |
| 3 | Dutch Books for Language Models | Isaiah Andrews, Suproteem Sarkar | econ.GN | 2026-09-02 |
| #4 | DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation | Vasileios Baltatzis, Mert Inan, Connor Gillis +4 | cs.CL | 2026-09-02 |
| #5 | EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction | Yuling Shi, Zhensu Sun, Junsen Dong +3 | cs.CL | 2026-09-02 |
| #6 | ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding | Jitai Hao, Ke Yang, Qiang Huang +1 | cs.CV | 2026-09-02 |
| #7 | HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks | Jongkyung Shin, Minguk Jeon, Chanwoo Park +1 | cs.CL | 2026-09-02 |
| #8 | From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution | Yuzhang Luo, Chenpeng Wang, Jianhui Chen +1 | cs.CL | 2026-09-02 |
| #9 | Untangling the Mechanisms of Misleading Context in Medical Question Answering | Robin Linzmayer, Noémie Elhadad | cs.CL | 2026-09-02 |
| #10 | Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills | Jianlyu Chen, Yuyang Hu, Hongjin Qian +8 | cs.AI | 2026-09-02 |
| #11 | Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection | Max Nelson, Hanoz Bhathena, Aviral Joshi +1 | cs.IR | 2026-09-02 |
| #12 | Language Models Can Control Their Own Attention | Namgyu Ho, Huzama Ahmad, Woosung Koh +3 | cs.CL | 2026-09-02 |
| #13 | Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases | Bernard Muller, László Tóth, LaVonne Roberts | cs.CL | 2026-09-02 |
| #14 | CORAL: An LLM-Native Harness for Production Recommender Systems | Muhammad Rafay Azhar, Yuhang Zhou, Gilbert Jiang +7 | cs.CL | 2026-09-02 |
| #15 | Door-in-the-Face Requests and Refusal Behaviour in Large Language Models | Til Jordan | cs.AI | 2026-09-02 |
| #16 | Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers | Xu Zou, Jie Tang | cs.CL | 2026-09-02 |
| #17 | DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models | Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi +7 | cs.CL | 2026-09-02 |
| #18 | From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs | Urja Pawar, Rajitha Ramanayake, Owen O'Neill +4 | cs.CL | 2026-09-02 |
| #19 | oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions | Haoqiang Guo, Xuyi Chen, Bo Ke +5 | cs.CL | 2026-09-02 |
| #20 | WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities | Jiska Beuk, Gerasimos Spanakis | cs.CL | 2026-09-02 |
| #21 | Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting | Ron Begleiter, Katya Egert Berg, Gilad Saban +1 | cs.AI | 2026-09-02 |
| #22 | TaRA: Training-Aware Low-Rank Adaptation Initialization | Taehyeon Kim, Eunhyeok Park | cs.CL | 2026-09-02 |
| #23 | Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction | Kenichi Fujita, Yusuke Ijima | cs.SD | 2026-09-02 |
| #24 | Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language | Vinmay Khandode, Sai Karthik Kosuri, Neil K. R. Sehgal +6 | cs.CL | 2026-09-02 |
| #25 | When Persona Attributes Improve Population Alignment in Large Language Models | Leon Fröhling, Jens Rupprecht, Markus Strohmaier +1 | cs.CL | 2026-09-02 |
| #26 | Debias-SparseGPT: Bias-Aware Pruning for Large Language Models | Irina Proskurina, Guillaume Metzler, Antoine Gourru +1 | cs.CL | 2026-09-02 |
| #27 | ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering | Adrien Mialland, Marc Plantevit, Julien Gallois +1 | cs.IR | 2026-09-02 |
| #28 | How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling | Katrin Rohrbacher, Björn Nieth, Emmanuelle Salin +2 | cs.CL | 2026-09-02 |
| #29 | PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation | Smitha Muthya Sudheendra, Jaideep Srivastava | cs.CL | 2026-09-02 |
| #30 | Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis | Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim +5 | cs.CL | 2026-09-02 |
| #31 | Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression | Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov | cs.LG | 2026-09-02 |
| #32 | When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models | Smitha Muthya Sudheendra, Jaideep Srivastava | cs.CL | 2026-09-02 |
| #33 | UTP-Bench: Uncertainty-aware Travel Planning Benchmark | Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick +3 | cs.AI | 2026-09-02 |
| #34 | Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking | Siyu Chen, Haoran Wang, Xiaojian Li +3 | cs.CL | 2026-09-02 |
| #35 | Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation | Egecan Çelik Evgin, İlknur Karadeniz, Olcay Taner Yıldız | cs.CL | 2026-09-02 |
| #36 | PolERo: Studying Political Evasion in Romanian | Gabriel Stefan, Sergiu Nisioi | cs.CL | 2026-09-02 |
| #37 | MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts | Matteo Greco, Anudeex Shetty, Andrea Tagarelli +1 | cs.CL | 2026-09-02 |
| #38 | NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning | Meixuan Chen, Hehan Li, Ruizhi Zhao +8 | cs.CL | 2026-09-02 |
| #39 | SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval | Zineb Lahrichi, Marc Ferras, Gaël Richard +1 | cs.SD | 2026-09-02 |
| #40 | SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning | Zhao Ji, Wenqing Chen, Zhixuan Chu +4 | cs.AI | 2026-09-02 |
| #41 | Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization | Bing Zheng, Zongyao Zhao, Wenming Yang | cs.IR | 2026-09-02 |
| #42 | DiffIE: Diffusion-based Open Information Extraction | Konstantin Fedorov, Valentin Malykh | cs.CL | 2026-09-02 |
| #43 | Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization | Bizhe Bai, Jiakang Yuan, Hongming Wu +8 | cs.CL | 2026-09-02 |
| #44 | Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds | Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6 | cs.AI | 2026-09-02 |
| #45 | SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology | Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko +2 | cs.AI | 2026-09-02 |
| #46 | Entangled Representations Amplify Collateral Damage in Unlearning | Evžen Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt | cs.LG | 2026-09-02 |
| #47 | Do Large Language Models Capture the Diversity in their Training Data? | Youqi Wu, Farzan Farnia | cs.CL | 2026-09-02 |
| #48 | CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging | Mingjie Zheng, Zihao Chen, Wenqing Chen +4 | cs.AI | 2026-09-02 |
| #49 | PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation | Yunhao Liu, Hong Phuc Pham, Jaehong Yoon | cs.CL | 2026-09-02 |
| #50 | From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X | Zhiyang Qi, Kazuhiro Ito, Jinghui Chen +5 | cs.SI | 2026-09-02 |
| #51 | APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering | Jie Ding, Rui Sun, Xinyuan Zhang +2 | cs.AI | 2026-09-02 |
| #52 | RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution | Zijian Zhao, Sen Li, Xialiang Tong +1 | cs.MA | 2026-09-02 |
| #53 | LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images | Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov | cs.CV | 2026-09-02 |
| #54 | Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search | Shiliang Xiao, Jingsong Wei, Yuzhi Liang +3 | cs.CL | 2026-09-02 |
| #55 | Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation | Ziqi Zhang, Emmanuele Chersoni, Mohammad Momenian | cs.CL | 2026-09-02 |
| #56 | OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction | Jaehoon Jeong, Jay-Yoon Lee | cs.CL | 2026-09-02 |
| #57 | A Layered Taxonomy for Chinese Learner Grammatical Error Annotation | Mengyang Qiu, Jungyeul Park | cs.CL | 2026-09-02 |
| #58 | EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision | Ziyuan Jin, Yuxuan Ge, Zheng Tian | cs.AI | 2026-09-02 |
| #59 | C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees | S M Rafiuddin, Atriya Sen | cs.CL | 2026-09-02 |
| #60 | AI agents reshape consensus formation in human groups | Lin Chen, Ziyi Liu, Xia Hu +1 | cs.CL | 2026-09-02 |
| #61 | text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation | Ritesh Kumar | cs.CL | 2026-09-02 |
| #62 | Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models | Haobo Xu, Sirui Chen, Yuanchen Bei +5 | cs.CL | 2026-09-02 |
| #63 | MASkills: Continual Skills Optimization for Multi-Agent LLM Systems | Huaiyuan Yao, Xiaoou Liu, Charles Fleming +2 | cs.AI | 2026-09-02 |
| #64 | Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage | Weifeng Jiang, Ruirui Chen, Qianren Mao +3 | cs.CL | 2026-09-02 |
| #65 | IDEEA: training-free Input-Dependent stEEring via Activation cluster matching | Zheng Wang, Muchen Li, Renjie Liao +1 | cs.CL | 2026-09-02 |
| #66 | XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression | Jundong Hu, Shekar Ramachandran | cs.LG | 2026-09-02 |
| #67 | Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models | Tianqi Xiao, Shiyao Cui, Minghao Zhang +2 | cs.MM | 2026-09-02 |
| #68 | HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs | Yihang Sun, Zhihan Zhu, Zhiyuan Jiang +3 | cs.CL | 2026-09-02 |
| #69 | Privacy Washing: Detecting Internal Contradictions in Privacy Policies | Thomas Brackin | cs.CY | 2026-09-02 |
| #70 | A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models | Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra | cs.CL | 2026-09-02 |
| #71 | The Dynamics of Continuous Mixture Collapse in Language Models | Ali Backour | cs.LG | 2026-09-02 |
| #72 | How Output Format Confounds Data Quality and Capability in Instruction Tuning | Chengguang Gan, Hanjun Wei, Yunhao Liang +3 | cs.CL | 2026-09-02 |
| #73 | Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation | Wenhui Chen, Zhifeng Li, Jie Zhou +5 | cs.LG | 2026-09-02 |
| #74 | NS-Copilot: An LLM-Driven Agent System for Autonomous Neuroscience Analysis | Wuche Liu, Yiran Qiao, Linlin Hou +4 | cs.CL | 2026-09-02 |
| #75 | Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens | Matteo He, William F. Shen, Xinchi Qiu +1 | cs.CL | 2026-09-01 |
| #76 | CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing | Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt +4 | cs.LG | 2026-09-01 |
| #77 | Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets | Kunal Jadhav, Siddhesh More | cs.CL | 2026-09-01 |
| #78 | Accurate in space, unreliable in time: how LLMs represent national cultural change | Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp | cs.CY | 2026-09-01 |
| #79 | GAPS: Dimension-Level Gates for Conditional Activation Steering | Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad +1 | cs.CL | 2026-09-01 |
| #80 | Thinking effort aligns between humans and reasoning models in abductive reasoning | Henry Arthur | cs.CL | 2026-09-01 |
| #81 | ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval | Aaryan Kapoor, Md Abdullah Al Hafiz Khan | cs.SE | 2026-09-01 |
| #82 | The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents | Jundong Hu, Shekar Ramachandran | cs.AI | 2026-09-01 |
| #83 | Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos | S M Masrur Ahmed, Jaspal Subhlok | cs.CL | 2026-09-01 |
| #84 | Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition | Weiming Li, Catarina Barata, Miguel Constante +1 | cs.CL | 2026-09-01 |
| #85 | Interpretable Symptom Vectors for Depression in a Large Language Model | Fangyi Zhu, Ajay Subramanian, Allison Constant +3 | cs.CL | 2026-09-01 |
| #86 | AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking | Chunggi Lee, Hanspeter Pfister | cs.CL | 2026-09-01 |
| #87 | TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding | Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari +3 | cs.CL | 2026-09-01 |
| #88 | How Do Prompt Variations Affect Energy Consumption in On-Device LLMs? | Wei Hu, Xiaolong Tu, Dawei Chen +3 | cs.CL | 2026-09-01 |
| #89 | Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization | Yixuan Wang, Freda Shi, Kanishka Misra | cs.CL | 2026-09-01 |
| #90 | VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages | Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda +4 | cs.CL | 2026-09-01 |
| #91 | MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models | Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5 | cs.CL | 2026-09-01 |
| #92 | When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic | Surya Saka | cs.AI | 2026-09-01 |
| #93 | SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition | Biraj Subedi | cs.CL | 2026-09-01 |
| #94 | Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives | Haibo Jin, Suijin Wang, Xucheng Yu +2 | cs.SE | 2026-09-01 |
| #95 | Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation | Himil Vasava, Ming Jiang | cs.CL | 2026-09-01 |
| #96 | Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation | Kefeng Duan, Dewu Zheng, Yanlin Wang +7 | cs.SE | 2026-09-01 |
| #97 | Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation | Kefeng Duan, Dewu Zheng, Yanlin Wang +8 | cs.SE | 2026-09-01 |
| #98 | CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses? | Damien Sileo, Dimitri Kachler | cs.CL | 2026-09-01 |
| #99 | The Rise of Verbal Reinforcement Learning | Kshitij Tayal, Arun Sharma, Genta Indra Winata +2 | cs.CL | 2026-09-01 |
| #100 | StudentSim: Training LLM-based Student Simulators | Ke Yang, Chenglong Wang, Michel Galley +4 | cs.CL | 2026-09-01 |
| #101 | Designing Proactive Thought Partners for Writing | Chao Zhang, Abe Davis, Chih-Wei Chen +1 | cs.HC | 2026-09-01 |
| #102 | The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally | Jundong Hu, Shekar Ramachandran | cs.LG | 2026-09-01 |
| #103 | Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics | Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin +3 | cs.CL | 2026-09-01 |
| #104 | Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs | Jingtan Wang, Arun Verma, Xiaoqiang Lin +4 | cs.CL | 2026-09-01 |
| #105 | From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix | Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov +11 | cs.CL | 2026-09-01 |
| #106 | Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers | Giovanni Bonetta, Matteo Merler, Davide Zago +2 | cs.AI | 2026-09-01 |
| #107 | From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification | Manish Gupta, Chaitanya Giri, Jayasimha Talur | cs.CL | 2026-09-01 |
| #108 | A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains | Ema Salkić, Alexander Fichtl, Philipp Ulrich +3 | cs.CL | 2026-09-01 |
| #109 | SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue | Stephanie Fong, Yiwen Jiang, Zimu Wang +12 | cs.CL | 2026-09-01 |
| #110 | Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall | Jacqueline He, Howard Yen, Shuyue Stella Li +9 | cs.CL | 2026-09-01 |
| #111 | GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions | Elias Stengel-Eskin, Newton Sander, Carlos Bonetti +4 | cs.CL | 2026-09-01 |
| #112 | AutoConcept: Training-Free Concept-Guided Reranking for Metadata-Available Composed Image Retrieval | Tianyu Wang, Tianjiao Wu | cs.IR | 2026-09-01 |
| #113 | HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? | Yuhao Wu, Jingyuan Zhang, Jiajun Shi +16 | cs.SE | 2026-09-01 |
| #114 | Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation | Yixuan Liu, Lin Chen, Zhuoqi Liu +2 | cs.DL | 2026-09-01 |
| #115 | From Rollouts to Recipes: Self-Contained Post-Training for LLMs | Yifei Li, Lingling Zhang, Muye Huang +3 | cs.CL | 2026-09-01 |
| #116 | EdiTikZ: Scientific Figure Editing from Revision Trajectories | Christian Greisinger, Zhixue Zhao, Steffen Eger | cs.AI | 2026-09-01 |
| #117 | When Tokenization is Secretly Output Supervision | Tanja Baeumel, Josef van Genabith, Simon Ostermann | cs.CL | 2026-09-01 |
| #118 | Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA | Haruto Sato, Yuki Tanaka, Ren Nakamura +2 | cs.CL | 2026-09-01 |
| #119 | InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations | Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3 | cs.CL | 2026-09-01 |
| #120 | Polish ModernBERT: The Long and Short of Polish Language Understanding | Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata +1 | cs.CL | 2026-09-01 |
| #121 | IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals | Md. Atabuzzaman, Christian Alexander, Chris Thomas | cs.CV | 2026-09-01 |
| #122 | Behaviorally Effective LoRA Writes Are Sparse and Structured | Haruto Sato, Yuki Tanaka, Ren Nakamura +2 | cs.CL | 2026-09-01 |
| #123 | How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation | Elitsa Yotkova, Violeta Kastreva, Petar Velkov +4 | cs.CL | 2026-09-01 |
| #124 | Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA | Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto | cs.CL | 2026-09-01 |
| #125 | Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs | Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2 | cs.CL | 2026-09-01 |
| #126 | Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR | Esther Xin | cs.CL | 2026-09-01 |
| #127 | CHARM: Character Hallucination for Multicultural Role Play Benchmark | Sunkyung Han, Nahyeon Park, Gaeun Seo +2 | cs.CL | 2026-09-01 |
| #128 | Probing Factual Knowledge Transfer with Training Data Interventions | Romina Oji, Marc Braun, Marcel Bollmann +2 | cs.CL | 2026-09-01 |
| #129 | VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models | Zhiqi Huang, Vivek Datla, Zhichao Xu +3 | cs.CL | 2026-09-01 |
| #130 | Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment | Mian Zhong, Katherine A. Keith, Anjalie Field | cs.CL | 2026-09-01 |
| #131 | Reliability Challenges in Diffusion Vision-Language Models | Md. Atabuzzaman, Chris Thomas | cs.CV | 2026-09-01 |
| #132 | MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval | Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro +2 | cs.IR | 2026-09-01 |
| #133 | Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents | Ruochen Zhou, Zhengyu Chen, Luan Zhang +3 | cs.CL | 2026-09-01 |
| #134 | Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models | Tian Fang, Gaël Guibon, Davide Buscaldi | cs.CL | 2026-09-01 |
| #135 | From Base Rollouts to RL Reasoning: A Budgeted Search Perspective | Wenhe Sun, Cunxiang Wang, Zijun Yao +1 | cs.CL | 2026-09-01 |
| #136 | What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal | Radin Shayanfar, Keheliya Gallaba, Ahmed E. Hassan | cs.SE | 2026-09-01 |
| #137 | Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation | Thibaut Thonet, Jos Rozen, Laurent Besacier | cs.CL | 2026-09-01 |
| #138 | Post-Training Science for Supervised Fine-Tuning | Charles O'Neill, Mudith Jayasekara, Harry Partridge | cs.LG | 2026-09-01 |
| #139 | Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment | Yin Fang, Qiao Jin, Shubo Tian +24 | cs.CL | 2026-09-01 |
| #140 | FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue | Hangyeul Lee, Juyoung Oh, Jaeyong Ko +5 | cs.AI | 2026-09-01 |
| #141 | CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs | Chaohui Guo, Michel Klein, Zhisheng Huang | cs.CL | 2026-09-01 |
| #142 | PersuaRL: Reinforcement Learning-Driven Multi-Expert Selection for Persuasive Dialogue Generation in Insurance | Rohan Kirti, Akash Ghosh, Aryan Vats +5 | cs.CL | 2026-09-01 |
| #143 | LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs | Muhammed Saeed, Simon Razniewski | cs.CL | 2026-09-01 |
| #144 | Subword Segmental BabyLMs: Learning to Tokenise for Sample-Efficient Pretraining | Francois Meyer | cs.CL | 2026-09-01 |
| #145 | On the Design Fundamentals of Pixel Text Representation Learning | Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4 | cs.CV | 2026-09-01 |
| #146 | Does task decomposition improve automatic NLG evaluation? | Sebastian Steindl, Nikos Voskarides, Alberto Gasparin +1 | cs.CL | 2026-09-01 |
| #147 | Overfitting Mitigation via Singular Value Decomposition in Minimum Bayes Risk Decoding | Riza Setiawan Soetedjo, Yusuke Sakai, Hidetaka Kamigaito +2 | cs.CL | 2026-09-01 |
| #148 | Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs | Zhaoliang Chen, Jie Fu | cs.AI | 2026-09-01 |
| #149 | EDRAC: Benchmarking Arabic Dialect Reading Comprehension | Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15 | cs.CL | 2026-09-01 |
| #150 | ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues | Huimin Wang, Zhengyi Zhao, Yutian Zhao | cs.CL | 2026-09-01 |
| #151 | Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation | Will Badr | cs.SE | 2026-09-01 |
| #152 | When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP | Shota Sato, Hajime Kiyama, Tosho Hirasawa +1 | cs.CL | 2026-09-01 |
| #153 | Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts | Nikolaos Xiros, Dimitrios Damianos, Maria-Eleni Zoumpoulidi +3 | cs.CL | 2026-09-01 |
| #154 | StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions | Chao Gao, Haijiang Liu, Qiyuan Li +3 | cs.CL | 2026-09-01 |
| #155 | Post-hoc Alignment of LLM-judges to Human Judgment Distribution | Sebastian Steindl, Nikos Voskarides, Alberto Gasparin +1 | cs.CL | 2026-09-01 |
| #156 | OUTLETS: Output-Length Prediction from Speculative Decoding Backbones | Weihuang Wen, Yingying Liu, Yichuan Liu +5 | cs.CL | 2026-09-01 |
| #157 | WorldBench: Culturally Grounded Benchmark for Multilingual Agents | Leonardo Ranaldi, Sherrie Shen, Jushi Kai +1 | cs.AI | 2026-09-01 |
| #158 | Lagged Coupling: Internal Representations Become Readable Before They Become Causal | Xining Xun | cs.CL | 2026-09-01 |
| #159 | PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition | Ziyan Gan, Fangxin Liu, Chenyang Guan +10 | cs.CL | 2026-09-01 |
| #160 | Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech | Che Hyun Lee, Sangkwon Park, Donghun Kang +4 | cs.CL | 2026-09-01 |
| #161 | SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models | Shiyu Li, Zi-Yuan Hu, Shijia Huang +3 | cs.CV | 2026-09-01 |
| #162 | Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close | Rania Elbadry, Ahmed Heakl, Saeed Almheiri +12 | cs.CL | 2026-09-01 |
| #163 | Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages | Michele Ciletti | cs.CL | 2026-09-01 |
| #164 | Disclosure-Gated User Simulation for Companion-Agent Evaluation | Yao Liu, Yu He | cs.CL | 2026-09-01 |
| #165 | PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian | Mohammad Hossein Shalchian, Mostafa Amiri, Amir Mahdi Sadeghzadeh | cs.CL | 2026-09-01 |
| #166 | Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling | Kangjia Zhao, Jiajun Li, Haozhan Shen +8 | cs.CL | 2026-09-01 |
| #167 | From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding | Raul Ortega, José Manuel Gómez-Pérez | cs.CV | 2026-09-01 |
| #168 | A Dataset for Modeling Iterative Problem-Solving | Fagun Patel, Sang T. Truong, Duc Q. Nguyen +4 | cs.CL | 2026-09-01 |
| #169 | DualStake: Dual-Path Confidence Calibration in Deep Research Agents | Yinuo Xu, Yuwei Liang, Jianjie Cheng +4 | cs.CL | 2026-09-01 |
| #170 | Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO | Prakhar Gupta, Vaibhav Gupta | cs.CL | 2026-09-01 |
| #171 | VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences | Yiwen Jiang, Yang Deng, Stephanie Fong +9 | cs.AI | 2026-09-01 |
| #172 | RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation | Zhongru Chen, Yuan Wu, Yi Chang | cs.AI | 2026-09-01 |
| #173 | Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry | Shengfang Zhai, Leo Marchyok, Yuling Shi +4 | cs.CL | 2026-09-01 |
| #174 | The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence | Genpei Zhang | cs.CV | 2026-09-01 |
| #175 | MemoryWalker: Stop Training Agents on Contexts They Never Saw | Zinco J, Xunjie Zhu, Shen Huang +3 | cs.LG | 2026-09-01 |
| #176 | Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources | Ivan Decostanzi, Michele Ronco, Sergio Consoli +8 | cs.AI | 2026-09-01 |
| #177 | Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers | Hyeonseop Yoon, Jeong-Eun Park | cs.CL | 2026-09-01 |
| #178 | Replacing Training with Memory: Listwise Selection for Text-to-SQL | Yeonseok Jeong, Soyoung Yoon, Seongjun Lee +1 | cs.SE | 2026-09-01 |
| #179 | Dense Process Supervision for Search Agents via Fact Utility Estimation | Rongzhi Zhu, Xiangyu Liu, Yi Liu +7 | cs.CL | 2026-09-01 |
| #180 | TWIX: a Two-Stage Approach for End-To-End Named Entity Recognition and Relation Extraction | Marco Martinelli, Laura Menotti | cs.CL | 2026-09-01 |
| #181 | Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents | Haoyang Chen, Yi Liu, Jianzhi Shao +3 | cs.AI | 2026-09-01 |
| #182 | Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online | Elisabeth Steffen, Helena Mihaljević | cs.HC | 2026-09-01 |
| #183 | TEIDAN: A Multilingual Multiparty Dialogue Corpus | Taiga Mori, Koji Inoue, Mikey Elmers +2 | cs.CL | 2026-09-01 |
| #184 | SFAD: Speculative Factuality-Aware Decoding | Guanqiao Chen, Di Wang, Lijie Hu | cs.CL | 2026-09-01 |
| #185 | Instella-MoE Technical Report | Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10 | cs.CL | 2026-09-01 |
| #186 | When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection | Utsab Ghosh, Roshni Chakraborty | cs.AI | 2026-09-01 |
| #187 | A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals | Yuri Son, Seunghee Kim, Hyuhng Joon Kim +1 | cs.CL | 2026-09-01 |
| #188 | Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning | Jinhu Qi, Minda Hu, Wentao Zhang +4 | cs.CL | 2026-09-01 |
| #189 | Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding | Chengguang Gan, Yunhao Liang, Hanjun Wei +2 | cs.CL | 2026-09-01 |
| #190 | How Do Language Models Choose Between Context and Memory? | Benjamin Shih, John Winnicki, Arianna Cao | cs.LG | 2026-09-01 |
| #191 | Measuring Optimal Transport in Transformer Depth | Alexandre Quemy | cs.CL | 2026-09-01 |
| #192 | Can Large Language Models Forecast What Researchers Study Next? | Fenghai Li, Zihan Tang, Haofei Yu +2 | cs.CL | 2026-09-01 |
| #193 | ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything | Yufan Dang, Shu Yao, Bowen Lai +6 | cs.AI | 2026-09-01 |
| #194 | Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports | Safayat Bin Hakim, Houbing Herbert Song | cs.CR | 2026-09-01 |
| #195 | Controllable Image Captioning with Prompt-Conditioned Scene Rewards | Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim | cs.CV | 2026-09-01 |
| #196 | A Certificate-Producing Cascade for Equational Implication: The SAIR EQT2 Stage 2 Solver | Haobo Ma, Wenlin Zhang, Manuel Israel Cázares | cs.CL | 2026-09-01 |
| #197 | Value Over Language Model: Detecting Original Contribution in Writing | Vibhhu Sharma, Thorsten Joachims, Sarah Dean | cs.AI | 2026-09-01 |
| #198 | SCoNE: Selective Context-aware Neuron Editing for Robust Retrieval-Augmented Generation | Chaewon Kim, Seo Yeon Park | cs.CL | 2026-09-01 |
| #199 | Visual Framing for News Stance Detection via Image Generation | Dahyun Lee, Jiyoung Han, Kunwoo Park | cs.CL | 2026-09-01 |
| #200 | Creative Generation via Multi-Agent Debate: Does Debate Suppress Diversity? | Tien Anh Nguyen, Khanh-Binh Nguyen, Van Dai Do +2 | cs.CL | 2026-09-01 |