PaperScope
LIVE · 2026-10-01 05:40 UTC

ModalFidelity: Routing Modalities for Deepfake Detection on a Budget

Oguzhan Baser, Kaan Kale, Sriram Vishwanath, Sandeep Chinchali

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38246 v1
Category
Submitted
2026-09-29

Abstract

Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.

Comment: 5 pages, 4 figures, 1 table. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch