PaperScope
LIVE · 2026-09-29 05:40 UTC

Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs

Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32333 v1
Category
Submitted
2026-09-26

Abstract

Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.

arXiv abs page · PDF · same-day batch