PaperScope
LIVE · 2026-09-25 05:40 UTC

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.29766 v1
Category
Submitted
2026-09-24

Abstract

Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.

Comment: Submitted to IEEE ICASSP 2027

arXiv abs page · PDF · same-day batch