PaperScope
LIVE · 2026-09-29 05:40 UTC

MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation

Ruihong Zeng, Jonathan Tonglet, Preslav Nakov, Iryna Gurevych

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33441 v1
Category
Submitted
2026-09-27

Abstract

Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers often check: whether an image's content is consistent with the context implied by its accompanying claim. To address this gap, we introduce MIC (Multimodal Inconsistency Checking), an AFC framework that assists human fact-checkers by detecting AI-generated multimodal misinformation and explaining inconsistencies using world knowledge. MIC first uses supervised fine-tuning (SFT) for task adaptation and then applies Group Relative Policy Optimization (GRPO) to directly optimize component-level verifiable rewards for verdict prediction, inconsistency type classification, visual evidence description, and world-knowledge explanation. We further introduce MIC-Bench, a benchmark comprising 8,812 image-claim instances derived from 4,406 claims, where each claim is paired with an authentic image and an AI-generated counterpart that introduces a controlled contextual inconsistency. Compared with SFT alone, GRPO further improves Macro-F1 by 4.67 and 4.11 points in the in-distribution and out-of-distribution settings, respectively, while also improving the semantic similarity of visual evidence descriptions and world-knowledge explanations to reference annotations. Our code and data are available at https://github.com/UKPLab/arxiv2026-mic.

Comment: Preprint under review

arXiv abs page · PDF · same-day batch