PaperScope
LIVE · 2026-09-29 05:40 UTC

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Peng Jun, Yiyi Zhou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34581 v1
Category
Submitted
2026-09-28

Abstract

Temporal video grounding is a key capability of advanced \emph{Multimodal Large Language Models} (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of \emph{On-policy distillation} (OPD) and propose a new training regime for MLLMs termed \emph{Counterfactual Attention Policy Distillation} (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0\% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.

arXiv abs page · PDF · same-day batch