PaperScope
LIVE · 2026-09-17 05:40 UTC

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16255 v1
Category
Submitted
2026-09-14

Abstract

We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-parameter model using only $\sim$900 uncertainty-selected examples, each augmented with synthetic chain-of-thought (CoT) rationales generated by a 4B teacher. Despite its minimal compute cost - under two hours on a single A100 GPU - our method enables the 2B model to outperform VLMs up to 4$\times$ larger, and generalize across CinePile, ActivityNet-QA, and MLVU, approaching the performance of its own 4B teacher. A key finding is that placing CoT rationales after the answer - contrary to standard prompting - substantially improves reasoning in compact models. This insight challenges prevailing CoT conventions and reveals new alignment strategies under limited model capacity. Our findings offer a practical blueprint for training deployable, reasoning-rich VLMs suited for mobile and edge applications.

Comment: 14 pages, 2 figures, 5 tables. Published in MultiMedia Modeling (MMM 2026), LNCS 16412

Journal: MultiMedia Modeling (MMM 2026), Lecture Notes in Computer Science, vol. 16412, pp. 567-580, Springer, Singapore, 2026

arXiv abs page · PDF · same-day batch