PaperScope
LIVE · 2026-09-09 05:40 UTC

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

Logesh Kumar Umapathi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07154 v1
Category
Submitted
2026-09-07

Abstract

We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the <=2B parameter division with 0.8279 on the held-out test set. Our system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass; It is obtained by distilling the junior perception module of a tool-using agentic pipeline, not the agent itself into a small student, using teacher traces filtered to those that answered correctly. it reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters. This raises a 27.1% base model to 81.4% on our held-out questions. The 2B backbone has 2.2132B parameters and therefore over the divisional limit, to make the entry admissable we prune the multilingual embedding table from 248,320 to 143,469 rows, reaching 1.9985B with provably identical logits on retained rows.

Comment: Winning solution technical report for the EgoLongQA track of the ECCV 2026 Wearable AI Grand Challenge

arXiv abs page · PDF · same-day batch