PaperScope
LIVE · 2026-09-29 05:40 UTC

You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

Yian Wang, Ali Ebrahimpour-Boroojeny, Hari Sundaram, Varun Chandrasekaran

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34137 v1
Category
Submitted
2026-09-28

Abstract

Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $κ$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.

Comment: 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

arXiv abs page · PDF · same-day batch