PaperScope
LIVE · 2026-09-29 05:40 UTC

DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time

Ziqi Ma, Hongqiao Chen, Georgia Gkioxari

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35704 v1
Category
Submitted
2026-09-28

Abstract

Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/

Comment: Project website: https://glab-caltech.github.io/dynatokens/

arXiv abs page · PDF · same-day batch