PaperScope
LIVE · 2026-09-29 05:40 UTC

CoViST: Visual Token Compression via Composable States

Qi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33397 v1
Category
Submitted
2026-09-27

Abstract

Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9\%, 99.5\%, and 98.1\% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8\%, 99.9\%, and 99.1\% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.

arXiv abs page · PDF · same-day batch