PaperScope
LIVE · 2026-10-06 05:40 UTC

Fitting Vision Adapters at Frontier Scales

Jaehoon Lee, Harry Partridge, Mudith Jayasekara, Charles O'Neill, Max Kirkby, Michael Psenka

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05897 v1
Category
Submitted
2026-10-05

Abstract

Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.

Comment: NeurIPS 2026 Workshop: Grounded and Faithful Vision-Language Models for Real-World Deployment

arXiv abs page · PDF · same-day batch