PaperScope
LIVE · 2026-09-29 05:40 UTC

UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation

Youssef Mansour, Enis Simsar, Fadime Sener, Markos Georgopoulos, Albert Pumarola, Ali Thabet, Edgar Schoenfeld

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32518 v1
Category
Submitted
2026-09-26

Abstract

Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In this work we present UnStep, a training-free wrapper that accelerates few-step causal video models at inference by running them with fewer diffusion transformer (DiT) steps than during distillation and limiting the temporal window retained in the attention KV cache. We propose two inference-only mechanisms to recover quality lost by step reduction and attention windowing: renoising the generated latent frames to a near-clean level and reusing the existing clean-cache pass to refine them, and applying truncated SVD to the DiT attention value and output projections. We also accelerate inference with a quality-preserving runtime stack for the DiT and VAE decoder, including more efficient attention calls and KV indexing, fused Triton RoPE with cached coefficients, and VAE decoding with optimized memory layout, precision, and convolution kernels. By reducing computation and optimizing the runtime stack, UnStep sets a new throughput regime for causal video diffusion, by running substantially faster than current methods, reaching 50 FPS on a single H100 without quality loss, and 77 FPS on GB200, all without retraining.

arXiv abs page · PDF · same-day batch