PaperScope
LIVE · 2026-09-29 05:40 UTC

RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models

Weili Zeng, Feng Tian, Shengqi Liu, Yichao Yan

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33217 v1
Category
Submitted
2026-09-27

Abstract

Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the generator's evolving computation and uses it to guide generation. Specifically, a separate timestep-free encoder is trained, through a timestep-conditioned predictor, to recover generator states across depths and noise levels from masked clean images. By excluding the reference noise and masked-out content from the encoder's input, we encourage the encoder to distill visual information that is recoverable from the visible image context and predictive of generator states. The representation learned from the generator's evolving states is then fed back to guide and improve the generator, whose updated states provide supervision for further representation learning, forming a reciprocal learning process. This reciprocal process improves multi-step generation and representation quality, as measured by frozen linear probing on ImageNet with latent-space SiT and pixel-space JiT, without an externally pretrained representation teacher. Across the two unconditional settings, FID decreases by $19.2$--$40.8\%$ relative to native training, while linear-probe accuracy improves by 6.80--10.04 percentage points over the best searched raw generator features. The learned representation also serves as a distributional metric for one-step JiT post-training, extending its role from instance-level alignment to distribution-level supervision.

arXiv abs page · PDF · same-day batch