PaperScope
LIVE · 2026-09-03 05:40 UTC

Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling

Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29291 v3
Category
Submitted
2026-08-29

Abstract

Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.

arXiv abs page · PDF · same-day batch