PaperScope
LIVE · 2026-09-30 05:40 UTC

Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37230 v1
Category
Submitted
2026-09-29

Abstract

Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.

Comment: 27 pages, 10 figures. Code available at https://github.com/LABA-SNU/seeing-is-not-addressing

arXiv abs page · PDF · same-day batch