PaperScope
LIVE · 2026-10-06 05:40 UTC

Diffusion Transformers are Provably Optimal In-context Generators

Guoji Fu, Tomoya Wakayama, Ryotaro Kawata, Atsushi Nitanda, Wee Sun Lee, Taiji Suzuki

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05333 v1
Submitted
2026-10-04

Abstract

Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the challenge is how to learn and sample from an output distribution that reflects this task uncertainty. In this work, we theoretically analyze how a Diffusion Transformer (DiT), pretrained across diverse tasks, learns and generates predictive distributions for a new query from demonstrations. We first show that the natural target to generate from finite demonstrations is not an output derived from estimating a single task, but rather a predictive distribution that captures the task uncertainty remaining after observing the demonstrations. We then prove that a DiT can learn this predictive distribution through score estimation, using attention to aggregate information from demonstrations and diffusion to generate samples. Owing to this property, with sufficient pretraining resources and diffusion sampling steps, the resulting DiT achieves the minimax optimal rate over a Hölder class of test-time tasks. These results imply that DiT acts as a statistically grounded in-context generator capable of generating distributions adapted to new tasks while retaining the uncertainty inherent in finite demonstrations.

arXiv abs page · PDF · same-day batch