PaperScope
LIVE · 2026-09-22 05:40 UTC

VISTA: Video-Injected Stylized Text-to-Animation

Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23817 v1
Category
Submitted
2026-09-20

Abstract

We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protocol with latent cycle consistency enables joint learning across separate semantically rich and stylistically diverse datasets. As a proof-of-concept for controllable animation synthesis, we validate VISTA on rendered motion-capture references: it achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content--style balance at inference time.

Comment: 3 pages, 1 figure

arXiv abs page · PDF · same-day batch