PaperScope
LIVE · 2026-10-06 05:40 UTC

Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood

Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06691 v1
Category
Submitted
2026-10-05

Abstract

Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.

Comment: 5 pages. Published in Interspeech 2026

Journal: Proc. Interspeech 2026, pp. 393-397

arXiv abs page · PDF · same-day batch