PaperScope
LIVE · 2026-09-30 05:40 UTC

Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu, Jinbao Xue, Ke Zhang, Junshi Huang, Kan Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36828 v1
Category
Submitted
2026-09-29

Abstract

Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.

Comment: preprint

arXiv abs page · PDF · same-day batch