PaperScope
LIVE · 2026-09-29 05:40 UTC

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

Zhiqiang Xia, Yang Li, Xinyuan Zhang, Yuchen Liu, Haoyu Lu, Jiaming Xu, Runyu Shi, Ying Huang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34942 v1
Category
Submitted
2026-09-28

Abstract

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.

Comment: 21 pages including references and appendix

arXiv abs page · PDF · same-day batch