PaperScope
LIVE · 2026-10-06 05:40 UTC

Task-Aware Joint Pruning and Distillation for Efficient Audio Deepfake Detection

Miao He, Peng Cheng, Zhongjie Ba, Qing Wen, Li Lu, Xin Yang, Kui Ren

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05264 v1
Category
Submitted
2026-10-04

Abstract

Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parameters) prevent deployment on resource-constrained devices. Existing compression methods, designed mainly for content-centric tasks, struggle to maintain competitive performance when directly adapted to deepfake detection. We propose a Task-Aware Joint Pruning and Distillation framework that combines cross-domain knowledge distillation with movement-guided structured pruning to transfer forgery-discriminative knowledge and preserve critical structures under aggressive compression. Our framework reduces the model to 31.9M parameters with 6.3$\times$ FLOPs reduction, with an average performance drop of only 1.30\% across multiple datasets compared to the uncompressed baseline, demonstrating strong potential for on-device deployment.

Comment: 6 pages, 4 figures, accepted to Interspeech 2026

arXiv abs page · PDF · same-day batch