PaperScope
LIVE · 2026-09-29 05:40 UTC

In-Token Learning for High-Fidelity Image Restoration via Diffusion Transformers

Xingfu Yi, Xiaoxue Yu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33523 v1
Category
Submitted
2026-09-27

Abstract

We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD ($2560{\times}1440$) inference and a tiled $12$K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.

Comment: Technical report, 25 pages, 12 figures. Preserves the earlier broader study underlying Fill2SR (ECCV 2026), including automatic colorization; Fill2SR subsequently developed the real-world super-resolution direction

arXiv abs page · PDF · same-day batch