PaperScope
LIVE · 2026-10-09 05:40 UTC

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.11775 v1
Category
Submitted
2026-10-08

Abstract

Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ${\sim}65\%$ higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.

Comment: 33 pages (12 non-appendix pages), 7 figures, published as a conference paper at ICML 2026

arXiv abs page · PDF · same-day batch