PaperScope
LIVE · 2026-09-10 05:40 UTC

Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators

Xinyu Chen, Adnan Mahmood, Mark Dras

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.09895 v1
Category
Submitted
2026-09-09

Abstract

Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliability difficult to compare. We introduce VidHalLoc, a benchmark that evaluates hallucination detection methods under a unified diagnostic evaluation protocol using 2,000 adversarial hallucination samples across Video Question Answering and Video Captioning tasks, spanning Ontology and Dynamic hallucination categories. To construct VidHalLoc efficiently, we introduce VideoHALO, a Harness Engineering-informed multi-agent workflow that decomposes data construction into four executable stages supported by a memory system and a communication protocol. Evaluation of fifteen methods reveals that the four dedicated detectors peak at an Overall accuracy of only 34.63%, indicating limited reliability across video hallucination types [Dataset Repository: https://huggingface.co/datasets/wesfggfd/VidHalLoc].

Comment: 29 pages, including appendices

arXiv abs page · PDF · same-day batch