PaperScope
LIVE · 2026-10-02 05:40 UTC

Embedding Prediction Helps Image Generation

Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.02203 v1
Category
Submitted
2026-10-01

Abstract

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Comment: Project page: https://sihanxu.me/nepa-dit

arXiv abs page · PDF · same-day batch