PaperScope
LIVE · 2026-09-09 05:40 UTC

Comparing Self-Supervised and Domain-Invariant Features for Cross-Domain Voice Phishing Detection

Jeongmin Lee, Seung Yun, Minkyu Lee, Ran Han, Yoonkyu Woo, Jinxia Huang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07079 v1
Category
Submitted
2026-09-07

Abstract

Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to privacy constraints; when available, only a handful of samples exist, insufficient for fine-tuning; and lightweight acoustic-only detection is needed as an alternative to large self-supervised models. We compare domain-invariant prosodic features and self-supervised representations (HuBERT, wav2vec2.0) through cross-domain evaluation-training on scenario-based actor recordings and testing on authentic criminal calls. Domain-invariant prosodic features achieve 69.5% F1 zero-shot and 71.0% with 5-shot learning. HuBERT achieves highest performance (94.2% F1, 5-shot), while wav2vec2.0 exhibits a precision-oriented detection profile (90.2% F1 with 99.4% precision, 5-shot). These findings reveal fundamental trade-offs: domain-invariant features enable zero-shot deployment when no real data exists, while SSL methods achieve higher performance but require real samples and compute.

Comment: 5 pages, 2 figures. Accepted at INTERSPEECH 2026

arXiv abs page · PDF · same-day batch