PaperScope
LIVE · 2026-09-30 05:40 UTC

BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36913 v1
Category
Submitted
2026-09-29

Abstract

Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.

Comment: 5 pages, 2 figures, 2 tables

arXiv abs page · PDF · same-day batch