PaperScope
LIVE · 2026-09-09 05:40 UTC

Feature Superposition in Neural Networks: From Theory to Practice

Dai Shi, Xiaoyu Li, Andi Han, José Miguel Hernández-Lobato

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06862 v1
Category
Submitted
2026-09-06

Abstract

Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then study how a network encodes those values in a lower-dimensional hidden representation. Empirical work, by contrast, seeks to identify the features encoded in trained networks and determine their role in computation. In this survey, we review the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions. To connect these theoretical accounts with evidence from trained networks, we compare practical methods for recovering and analyzing features and examine what their evaluations establish. Since accurate activation reconstruction alone does not establish feature identity or causal use, we discuss the methods' documented failures and applications in light of the evidence available for these different claims. Finally, we assess previously stated open problems and identify remaining theoretical and empirical questions about superposition in trained networks. We hope our work can pave the way for a deeper understanding of superposition and more reliable methods for interpreting neural networks.

arXiv abs page · PDF · same-day batch