PaperScope
LIVE · 2026-09-03 05:40 UTC

Denoising Diffusion Generative Models Secretly Calculate Attentions

Farzan Haddadi, Leila Monfared, Ebrahim Rezaii, Mohammadreza Malek-Mohammadi, Pejman Zakalvand, Narges Mokhtari

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00885 v1
Submitted
2026-09-01

Abstract

Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.

Comment: submitted to IEEE Trans on Pattern Recog. Machine Intellig

arXiv abs page · PDF · same-day batch