PaperScope
LIVE · 2026-09-29 05:40 UTC

Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation

Xiaoyu Ma, Chen Yang, Hao Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32188 v1
Category
Submitted
2026-09-26

Abstract

Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.

Comment: submitted to ICLR 2027

arXiv abs page · PDF · same-day batch