PaperScope
LIVE · 2026-09-23 05:40 UTC

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

Robert Sutherland, Stefan Goetze, Jon Barker

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.25948 v1
Category
Submitted
2026-09-22

Abstract

Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.

Comment: Accepted to the International Workshop on Acoustic Signal Enhancement (IWAENC), Cremona, Italy, September 2026

arXiv abs page · PDF · same-day batch