| 1 | Do speech foundation models really learn words? | Robin Huo, Ewan Dunbar | cs.CL | 2026-09-09 |
| 2 | Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition | Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin +1 | eess.AS | 2026-09-09 |
| 3 | AVSRBench: A Multi-Condition AVSR Benchmark | Rishabh Jain, Naomi Harte | eess.AS | 2026-09-09 |
| #4 | The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding | Gilad D. Landau, Dulhan Jayalath, Oiwi Parker Jones | cs.CL | 2026-09-09 |
| #5 | Politics of Feelings: Emotional Expression and Legislative Effectiveness in the U.S. Congress | Segun Aroyehun | cs.CL | 2026-09-09 |
| #6 | NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments | Niramay M. Patel, Bibek Behera, Raksha Sharma | cs.SD | 2026-09-09 |
| #7 | Orukeet: Multilingual ASR with Frozen Gabor Kernels | Nathan Roll, Irene Yi, Büşra Marşan +6 | cs.SD | 2026-09-09 |
| #8 | Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations | Soumyadeep Roy | cs.SD | 2026-09-09 |
| #9 | Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS | Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1 | cs.SD | 2026-09-09 |
| #10 | Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services | Yonghyun Jun, Jimin Lee, Hwan Chang +3 | cs.CL | 2026-09-09 |
| #11 | Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields | Cy Gorman, Yihang Yao | cs.AI | 2026-09-09 |
| #12 | $S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants | Heyang Liu, Jiayi Huang, Wenyang Xiao +8 | cs.CL | 2026-09-09 |
| #13 | MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short | Kristin Gnadt, Maximilian Meidinger, Matthias Aßenmacher | cs.CL | 2026-09-09 |
| #14 | Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis | Hong Nguyen, Sean Foley, Christina Hagedorn +4 | cs.SD | 2026-09-09 |
| #15 | StreamAlign: Streaming Text-Aligned Speech Tokenization | Kang-wook Kim, Jinyoung Park, Jinsoo Kim +3 | cs.CL | 2026-09-09 |
| #16 | X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS | Zehan Liu, Carl Chen, Rime Wen +7 | cs.CL | 2026-09-09 |
| #17 | SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia | Jingyi Liao, Wenyu Zhang, Zhuohan Liu +6 | cs.CL | 2026-09-09 |
| #18 | Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference | Yaohan Guan, Yen-Ju Lu, Yuzhe Wang +5 | cs.MA | 2026-09-09 |
| #19 | BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models | Shivam Singh, Aditya Yadavalli, Catherine Arnett +1 | cs.CL | 2026-09-09 |
| #20 | The Mutations of Machine Speech | Mauricio Figueroa | cs.CL | 2026-09-08 |
| #21 | Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models | Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1 | cs.SD | 2026-09-08 |
| #22 | Omni Interaction Agent Technical Report | Orantqing, Shengpeng Ji, Junlong Tong +20 | eess.AS | 2026-09-08 |
| #23 | AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing | Ziyang Ma, Zhikang Niu, Wenming Tu +30 | cs.SD | 2026-09-08 |
| #24 | From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection | Mengzhe Geng, Yujia Lu, Patrick Littell +2 | cs.SD | 2026-09-08 |
| #25 | TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context | Fritz Cremer, Jonathan Cremer | cs.SD | 2026-09-08 |
| #26 | X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR | Zhiwei Lin, Kaiqi Fu, Rime Wen +5 | cs.SD | 2026-09-08 |
| #27 | Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models | Luka Debevc, Nishan Chatterjee, Antoine Doucet +2 | cs.CL | 2026-09-08 |
| #28 | Detecting Authorship in Political Texts with Inductive Stylometry | Gennadii Iakovlev, Levente Littvay | cs.CL | 2026-09-08 |
| #29 | Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models | Bella Godiva, Yeonju Kim, Yong Man Ro | cs.SD | 2026-09-08 |
| #30 | EviSI: An Evaluation Agent for Simultaneous Interpreting | Ben Yan, Zongyao Li, Daimeng Wei +5 | cs.CL | 2026-09-08 |
| #31 | ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion | Richard Yucheng He, Baodong Cao, Chen Xu +2 | cs.CL | 2026-09-08 |
| #32 | Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech | Chibuzor Okocha, Christan Grant, Zoey Liu | cs.CL | 2026-09-07 |
| #33 | Syntactic Patterns and Stylistic Functions in Narrative Prose: A Rule-Based and Machine-Learning Approach | Stefana Janicijevic | cs.CL | 2026-09-07 |
| #34 | Qwen-Audio-3.0-ASR Technical Report | Chuanmeng Bian, Daren Chen, Peixin Chen +42 | cs.CL | 2026-09-07 |
| #35 | Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition | Peng Xie | cs.CL | 2026-09-07 |
| #36 | RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems | Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong +3 | cs.AI | 2026-09-07 |
| #37 | Iterative Audio Separation with Mixture Consistency via MIMO Model Extension | Yukara Ikemiya, WeiHsiang Liao, Yuki Mitsufuji | cs.SD | 2026-09-07 |
| #38 | KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction | Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji Ogawa | cs.SD | 2026-09-07 |
| #39 | AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation | Suah Choi, Tae-Young Lee, Gyeong-Moon Park | cs.CV | 2026-09-07 |
| #40 | When Speech Meets Lips: Interpretable Audio-Visual Synchronization for L2 Pronunciation Assessment | Bowen Yu, Mingyu Huang, Yishen Liu +1 | cs.CV | 2026-09-06 |
| #41 | MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games | Hwang Yechan, Bae Sangjun, Kim Jeongmo +2 | cs.AI | 2026-09-06 |
| #42 | Visual Search Augmented Chain-of-Thought Reasoning for Attribute Value Extraction from Product Videos | Tong Wu, Ming Cheng, Jiazhen Hu +2 | cs.CL | 2026-09-06 |
| #43 | PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation | Yuchen Sun, Qian Yang, Jun Wang +4 | cs.MM | 2026-09-04 |
| #44 | ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults | Weide Zhan, Qumu Shaqu, Yuanqing Liu +6 | cs.AI | 2026-09-04 |
| #45 | Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion | Xu Lin, Ke Wang, Hui Kang +1 | cs.CV | 2026-09-04 |
| #46 | Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks | Shivang Rawat, Mirko Morello, Flaviano Morone +1 | cs.LG | 2026-09-03 |
| #47 | Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis | Sanyuan Chen, Min-Jae Hwang, Sho Inoue +12 | cs.CL | 2026-09-03 |
| #48 | Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations | Yoto Fujita, Simon Leglaive, Laurent Girin | cs.SD | 2026-09-03 |
| #49 | Test-time adaptation for speech enhancement with an autoregressive speech prior | Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda +1 | cs.SD | 2026-09-03 |
| #50 | Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech | Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong +3 | cs.CL | 2026-09-03 |