PaperScope
LIVE · 2026-10-06 05:40 UTC

RoMod: Temporal Routing Modulation via Mixture-of-Experts for Video Anomaly Detection

Chao Huang, Pengfei Wei, Benfeng Wang, Chengliang Liu, Wei Wang, Li Shen, Wenqi Ren, Xiaochun Cao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05131 v1
Category
Submitted
2026-10-04

Abstract

Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make their internal computation easier to inspect. With a fully frozen backbone and no additional training, we find that anomaly-related evidence is concentrated in a small set of experts. These experts recur across layers, spontaneously specialize in different anomaly types, and together form a dynamic routing subnetwork. We further show that the output channels most strongly influenced by these experts are also the hidden dimensions that contain the most anomaly-relevant information. Routing statistics can therefore serve as an internal anomaly cue that complements semantic features.Based on these findings, we propose RoMod, an efficient VAD framework trained with only \(5\%\) of weakly labeled videos. RoMod includes a Routing-Modulated Fusion module, RoMF, and a Routing-aware Temporal Network, RoTN. RoMF uses routing signals to adaptively recalibrate hidden semantic channels. Its design also prevents the routing branch from bypassing semantic features and making predictions on its own. RoTN captures the temporal evolution of anomalies from onset to persistence and termination. Experiments on three benchmarks show that RoMod achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.

arXiv abs page · PDF · same-day batch