PaperScope
LIVE · 2026-09-21 05:40 UTC

DiaVLo: Diagnosing Behaviours of Vision-Language Models

Lorenzo Corti, Jie Yang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.22008 v1
Category
Submitted
2026-09-18

Abstract

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.

Comment: 34 pages. To appear in EMNLP 2026 (findings)

arXiv abs page · PDF · same-day batch