PaperScope
LIVE · 2026-09-17 05:40 UTC

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Minu Kim, Ji Sub Um, Hoirin Kim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16458 v1
Category
Submitted
2026-09-15

Abstract

Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.

Comment: Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch