PaperScope
LIVE · 2026-10-07 05:40 UTC

Quality-Aware Self-Correcting Speech Translation on an Edge Device

Zubair Ajmal Farooq, Diptesh Kanojia

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07545 v1
Category
Submitted
2026-10-06

Abstract

We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold $τ$. We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at $τ=0.90$ produces statistically significant improvements over greedy decoding on BLEU (+0.67, p<0.001), ChrF (+0.51, p<0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p>0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1$\to$M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.

Comment: 7 pages, 2 figures, 4 tables. Full paper submitted to the Convergence 2026 proceedings; poster presented at Convergence 2026, University of Surrey. Code: https://github.com/juebae/speech-translation_edge_device

arXiv abs page · PDF · same-day batch