PaperScope
LIVE · 2026-09-29 05:40 UTC

When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context

Yuxing Cheng, Yuan Wu, Yi Chang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34781 v1
Category
Submitted
2026-09-28

Abstract

Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.

arXiv abs page · PDF · same-day batch