PaperScope
LIVE · 2026-09-29 05:40 UTC

MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

Ashim Dahal, Bikramjit Banerjee

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33998 v1
Category
Submitted
2026-09-27

Abstract

Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of $8.9\%$ while slightly improving accuracy overall.

arXiv abs page · PDF · same-day batch