In-Token Learning for High-Fidelity Image Restoration via Diffusion Transformers
Xingfu Yi, Xiaoxue Yu
Abstract
We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD ($2560{\times}1440$) inference and a tiled $12$K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.