PaperScope
LIVE · 2026-10-01 05:40 UTC

Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization

Yehoshua Dissen, Joseph Keshet, Eduard Golshtein

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39162 v1
Category
Submitted
2026-09-30

Abstract

Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.

Comment: submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch