PaperScope
LIVE · 2026-09-30 05:40 UTC

VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

Zheng Jiang, Houde Qian, Yiming Chen, Ling Li, Chaoyang Li, Yueqi Li, Yuxuan Liu, Lifeng Sun

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38086 v1
Category
Submitted
2026-09-29

Abstract

Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.

arXiv abs page · PDF · same-day batch