PaperScope
LIVE · 2026-09-25 05:40 UTC

Learning a Flow to Self-Supervised Representations

Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.29350 v1
Submitted
2026-09-24

Abstract

Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.

Comment: 33 pages, 2 figures, including appendix

arXiv abs page · PDF · same-day batch