PaperScope
LIVE · 2026-09-29 05:40 UTC

SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models

Boyu Zhang, Yangming Cheng, Ning Zhang, Pengfei Liu, Weijie Li, Yifan Gao, Hangyu Li, Litong Gong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34525 v1
Category
Submitted
2026-09-28

Abstract

Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.

arXiv abs page · PDF · same-day batch