PaperScope
LIVE · 2026-10-06 05:40 UTC

MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06378 v1
Category
Submitted
2026-10-05

Abstract

The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.

Comment: 18 pages, 4 figures, 19 tables

arXiv abs page · PDF · same-day batch