PaperScope
LIVE · 2026-09-24 05:40 UTC

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

Sanghee Park, Kee-Eung Kim

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.27395 v1
Category
Submitted
2026-09-23

Abstract

Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.

Comment: Accepted to EMNLP 2026 Findings. 29 pages, 22 figures, 21 tables

arXiv abs page · PDF · same-day batch