PaperScope
LIVE · 2026-09-24 05:40 UTC

Hidden not Deleted: How Networks Suppress Entangled Features

Akash Samanta, Manish Pratap Singh, Debasis Chaudhuri

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.27593 v1
Category
Submitted
2026-09-23

Abstract

Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.

arXiv abs page · PDF · same-day batch