PaperScope
LIVE · 2026-09-28 05:40 UTC

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

Filip Tăşădan, Ema Tomanová, Ondrej Lopuch, Paweł Bilko, Anders Søgaard

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.30839 v1
Category
Submitted
2026-09-25

Abstract

Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.

Comment: Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch