PaperScope
LIVE · 2026-09-29 05:40 UTC

FuseAlign: Forced Alignment in the Wild

Mithilesh Vaidya, Stephen Bailey, Sumukh Badam, Matthew Bendel, Xingzhe He

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33650 v1
Category
Submitted
2026-09-27

Abstract

Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.

Comment: Preprint. Under review

arXiv abs page · PDF · same-day batch