PaperScope
LIVE · 2026-09-17 05:40 UTC

Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models

Ziheng Ren, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.17269 v1
Category
Submitted
2026-09-15

Abstract

Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible perception, and environmental decision-making, such hallucinations can create real-world safety risks. We propose Semantic-Spatial Agreement Verification (SSAV), a training-free method for verifying object claims. A visually grounded claim should remain stable across semantically equivalent queries and repeatedly localize to the same image region. SSAV aggregates multiple prompts to estimate semantic support and reduce sensitivity to query wording. Query-Induced Regional Verification (QIRV) combines cross-query region persistence, spatial overlap, and relative candidate dominance to identify isolated high responses and dispersed localizations. A geometric mean fuses semantic and spatial evidence, lowering the verification score when either branch lacks support. Experiments on three base models and multiple evaluation protocols show that SSAV effectively mitigates object hallucination. On LLaVA-1.5-7B, accuracy averaged across COCO, A-OKVQA, and GQA improves by 1.81 and 3.17 percentage points under POPE Popular and Adversarial, respectively, while CHAIRs decreases from 49.40% to 32.80%. These results show that cross-query semantic stability and regional consistency provide interpretable external visual evidence for object claims.

arXiv abs page · PDF · same-day batch