PaperScope
LIVE · 2026-09-29 05:40 UTC

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

Jiajun Xu, Menglu Li, Xiao-Ping Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33375 v1
Submitted
2026-09-27

Abstract

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.

Comment: 5 pages, 2 figures, 3 tables. Prepared for submission to ICASSP 2027

arXiv abs page · PDF · same-day batch