PaperScope
LIVE · 2026-09-03 05:40 UTC

FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

Shengyao Chen, Xianbing Sun, Liqing Zhang, Jianfu Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.30450 v1
Category
Submitted
2026-08-31

Abstract

Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose keypoints that frequently fail under large motions and occlusions, causing boundary artifacts and temporal inconsistency. A further limitation is that most approaches rely solely on attention mechanisms for temporal modeling, providing no explicit motion supervision. We propose FlowVVTON, a mask-free framework that eliminates parsing mask dependency entirely. Optical flow is used solely as a training-time supervision signal: a flow-warped latent loss, applied across all layers of the generation model, enforces multi-scale temporal consistency by aligning adjacent-frame features under explicit physical motion constraints. A two-stage training strategy establishes mask-free spatial alignment before introducing flow-guided temporal supervision. Experiments on TikTokDress show that FlowVVTON outperforms baselines by substantial margins, particularly in temporal consistency (5.7$\times$ VFID-R improvement over SwiftTry), while requiring no segmentation masks, pose keypoints, or region annotations at any stage.

arXiv abs page · PDF · same-day batch