PaperScope
LIVE · 2026-10-06 05:40 UTC

A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

Yilin Yang, Jun-Tao Tang, Kengyi Wang, Siyuan Su, Gaoyong Luo, Mingda Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05413 v1
Category
Submitted
2026-10-04

Abstract

Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.

Comment: Code is available at https://github.com/JuntaoTang/MLLM-VisionEncoder-Eval

arXiv abs page · PDF · same-day batch