PaperScope
LIVE · 2026-09-03 05:40 UTC

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.01888 v1
Category
Submitted
2026-09-01

Abstract

Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.

arXiv abs page · PDF · same-day batch