| 1 | TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context | Fritz Cremer, Jonathan Cremer | cs.SD | 2026-09-08 |
| 2 | Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values | Yuemei Xu, Kexin Xu, Jian Zhou +3 | cs.CL | 2026-09-08 |
| 3 | Compositional Multilingual and Behavioral Attribute Steering | Hyun Gu Kang, Daniil Gurgurov, Tanja Baeumel +2 | cs.CL | 2026-09-08 |
| #4 | Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning | Tejasvi C. Addagada | cs.CL | 2026-09-08 |
| #5 | EMBLEM: Enhancing Multi-script Table Detection through Masking | Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan | cs.LG | 2026-09-08 |
| #6 | Tracing Stereotypes from Representation to Output in Multilingual LLMs | Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen +2 | cs.CL | 2026-09-08 |
| #7 | IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA | Yuwen Chiu | cs.CL | 2026-09-08 |
| #8 | Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture | Hokky Situngkir | cs.AI | 2026-09-08 |
| #9 | Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions | Minh Duc Bui, Mario Sanz-Guerrero, Abteen Ebrahimi +4 | cs.CL | 2026-09-07 |
| #10 | Qwen-Audio-3.0-ASR Technical Report | Chuanmeng Bian, Daren Chen, Peixin Chen +42 | cs.CL | 2026-09-07 |
| #11 | Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy | Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos | cs.RO | 2026-09-07 |
| #12 | Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model | Logesh Kumar Umapathi | cs.CV | 2026-09-07 |
| #13 | LoGAN: Multilingual Font Localization with Generative Agents | Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein | cs.CV | 2026-09-07 |
| #14 | AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths | MaoXun Huang, Zhenxing Zhang, Claire Cardie | cs.CL | 2026-09-06 |
| #15 | A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation | Hari Krishnan K, Sudarsun Santhiappan | cs.CL | 2026-09-06 |
| #16 | Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus | Lifeng Han, Jiahui Liang, Anna Latusek +6 | cs.CL | 2026-09-06 |
| #17 | Discovering Translation-Worthy Languages with E-Values | Wajdi Ben Saad, Safa Madiouni | cs.CL | 2026-09-06 |
| #18 | Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space | Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala +1 | cs.CL | 2026-09-06 |
| #19 | EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages | Aleix Sant, Jordi Luque, Carlos Escolano | cs.CL | 2026-09-04 |
| #20 | MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models | Changming Xiao, Zhenliang Ni, Jinhui He +2 | cs.AI | 2026-09-04 |
| #21 | MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain | Sourav Malakar, Harshit Nigam, Akash Ghosh +4 | cs.CL | 2026-09-04 |
| #22 | A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures | Oskar Holmström, Marcel Bollmann, Marco Kuhlmann | cs.CL | 2026-09-04 |
| #23 | Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3 | Giang Son Nguyen, Nhi Ngoc-Yen Nguyen, Wray Buntine +1 | cs.CL | 2026-09-04 |
| #24 | Choosing the Right Language Mode at Inference Time for Multilingual Reliability | Ekata Mitra, Ameeta Agrawal | cs.CL | 2026-09-04 |
| #25 | Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation | Hasan Alkhder, Mohammad Abboush, Igor Tchappi +2 | cs.CL | 2026-09-03 |
| #26 | IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition | Nazim-E-Alam, Tarek Rahman, Md Kishor Morol | cs.CV | 2026-09-03 |
| #27 | A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval | Santiago Poveda-Gutiérrez, Hideki Nakayama, Mayumi Bono | cs.CV | 2026-09-03 |
| #28 | IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks | Saikat Mondal, Mamta, Deeksha Varshney +2 | cs.CL | 2026-09-03 |
| #29 | Typological Feature Prediction with Large Language Models: An In-Context Learning Approach | Qianwen Wang, York Hay Ng, Aditya Khan +1 | cs.CL | 2026-09-03 |
| #30 | Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations | Karthika Nhayakkat, Rajat Verma, Maharaj Brahma +4 | cs.CL | 2026-09-03 |
| #31 | Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases | Bernard Muller, László Tóth, LaVonne Roberts | cs.CL | 2026-09-02 |
| #32 | Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy | Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner +6 | cs.SI | 2026-09-02 |
| #33 | WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities | Jiska Beuk, Gerasimos Spanakis | cs.CL | 2026-09-02 |
| #34 | MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts | Matteo Greco, Anudeex Shetty, Andrea Tagarelli +1 | cs.CL | 2026-09-02 |
| #35 | VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages | Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda +4 | cs.CL | 2026-09-01 |
| #36 | MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models | Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5 | cs.CL | 2026-09-01 |
| #37 | Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs | Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2 | cs.CL | 2026-09-01 |
| #38 | Probing Factual Knowledge Transfer with Training Data Interventions | Romina Oji, Marc Braun, Marcel Bollmann +2 | cs.CL | 2026-09-01 |
| #39 | On the Design Fundamentals of Pixel Text Representation Learning | Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4 | cs.CV | 2026-09-01 |
| #40 | EDRAC: Benchmarking Arabic Dialect Reading Comprehension | Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15 | cs.CL | 2026-09-01 |
| #41 | WorldBench: Culturally Grounded Benchmark for Multilingual Agents | Leonardo Ranaldi, Sherrie Shen, Jushi Kai +1 | cs.AI | 2026-09-01 |
| #42 | TEIDAN: A Multilingual Multiparty Dialogue Corpus | Taiga Mori, Koji Inoue, Mikey Elmers +2 | cs.CL | 2026-09-01 |
| #43 | The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space | Jacob Brinton, Jannik Brinkmann, Mark Crovella +1 | cs.CL | 2026-09-01 |
| #44 | Latent Mechanisms of Language Control in Multilingual Language Models | Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly | cs.CL | 2026-08-31 |
| #45 | Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries | Phuong Anh Nguyen, Jill Noorily, Matthew Flathers +7 | cs.CY | 2026-08-31 |
| #46 | Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs | Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut | cs.CL | 2026-08-31 |
| #47 | Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models | Melina Morch, Daniel Braun | cs.CL | 2026-08-31 |
| #48 | Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature? | Nedjma Ousidhoum, Noopur Zambare, Mohamed Abdalla | cs.CL | 2026-08-31 |
| #49 | NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference | Aurélien Lac, Tony Wu | cs.IR | 2026-08-31 |
| #50 | Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages? | Donghoon Han, SungHyun Moon, Aidyn Zhakatayev +2 | cs.CL | 2026-08-31 |