PaperScope
LIVE · 2026-09-29 05:40 UTC

Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy

Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33935 v1
Category
Submitted
2026-09-27

Abstract

Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.

Comment: Project page: https://iandrover.github.io/video_analogy/

arXiv abs page · PDF · same-day batch