PaperScope
LIVE · 2026-09-30 05:40 UTC

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37090 v1
Category
Submitted
2026-09-29

Abstract

Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.

Comment: 13 pages. Submitted to IEEE Transactions on Mobile Computing

arXiv abs page · PDF · same-day batch