PaperScope
LIVE · 2026-10-01 05:40 UTC

TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning

Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39335 v1
Category
Submitted
2026-09-30

Abstract

Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to two main factors: (1) the insufficient utilization of rich garment reference information, and (2) the lack of explicit positional modeling between garment and video representations during cross-modal interaction, which weakens fine-grained local correspondence. To address these issues, we propose TexTailor, a high-fidelity video virtual try-on framework built upon a pretrained video Diffusion Transformer. Specifically, we introduce a timestep-adaptive modulation mechanism to dynamically adjust garment visual representations throughout denoising. We further develop a frame-aligned positional encoding strategy to strengthen garment-to-video correspondence, together with a multi-source injection design that reduces interference among heterogeneous conditions. Extensive experiments on multiple video virtual try-on benchmarks, including the high-resolution Eevee dataset, demonstrate that TexTailor achieves competitive performance in garment detail preservation, temporal consistency, and overall video quality.

arXiv abs page · PDF · same-day batch