PaperScope
LIVE · 2026-09-03 05:40 UTC

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

Yu Cheng, Arushi Goel, Hakan Bilen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29374 v1
Category
Submitted
2026-08-29

Abstract

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from tool failures. We propose ReVISE, a framework that equips MLLMs with verification and dynamic error recovery for tool-augmented reasoning. ReVISE introduces (1) a curated training dataset that supervises reflective behaviors, enabling models to validate tool-derived evidence, reformulate queries when visual mismatches arise, and fall back to intrinsic grounding when external tools are unreliable; and (2) a reinforcement learning based targeted rewards that encourage internal reflection and penalize spatial misalignment. Experiments on several benchmarks demonstrate consistent improvements over existing methods, highlighting the importance of error detection and correction in tool-augmented multimodal reasoning.

Comment: accepted by EMNLP2026

arXiv abs page · PDF · same-day batch