PaperScope
LIVE · 2026-09-09 05:40 UTC

Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection

Sébastien Thuau, Amira Gran, Siba Haidar, Rachid Chelouah

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07455 v1
Category
Submitted
2026-09-07

Abstract

How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained? Most weakly supervised VAD methods assume centralized training; recent VLM-based extensions further rely on dense inference, generated explanations, or additional adaptation. We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. We study two VLM feedback interfaces: parsed text-generation decisions and a logit-based interface that extracts a continuous anomaly score from next-token Yes/No probabilities. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-Instruct show that text-generation verification can improve frame-level AUC after diagnostic temporal post-processing, but remains sensitive to prompts, parsers, model choice, and smoothing. In contrast, the logit interface provides a fixed parser-free signal that improves both frame-level AUC and frame-level AP over the MIL baseline across both VLMs, without temporal post-processing in its main configuration. Since suspect segments are updated independently once available, next-token logit feedback provides a simple segment-local alternative to text-generation verification.

Comment: 6 pages, 1 figure, AVSS 2026

arXiv abs page · PDF · same-day batch