PaperScope
LIVE · 2026-10-06 05:40 UTC

Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

Paula A. Perez-Toro, Tomas Arias-Vergara, Annette Schwarz, Abner Hernandez, Andreas Horr, Cornelia Kristen, Andreas Maier

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06430 v1
Category
Submitted
2026-10-05

Abstract

Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions and speaker-independent read--spontaneous transfer. Whisper performs best under matched conditions, reaching 0.489 nine-way UAR and 148 km median geolocation error, but drops to 0.11/0.18 UAR across transfer directions and 363 km geolocation error. Self-supervised models show a similar degradation, whereas speaker embeddings are less discriminative in-domain but more robust under transfer. This contrast is consistent across classification and geolocation. Across representations, robustness is associated with how little a representation shifts between styles (style-invariance), for which crossstyle speaker retrieval is an interpretable proxy. Age, sex, sentence-overlap, and duration controls do not account for the gap, although channel characteristics contribute. These results show that strong matched-condition performance does not indicate robust regional information.

Comment: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

arXiv abs page · PDF · same-day batch