PaperScope
LIVE · 2026-09-03 05:40 UTC

Structure Aware Neural Architecture Search for Mixture of Experts

Petr Babkin, Oleg Bakhteev

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29817 v1
Category
Submitted
2026-08-30

Abstract

Neural Architecture Search (NAS) has so far rarely been applied to Mixture-of-Experts (MoE) models, and existing MoE designs leave the alignment between experts and the structure of the data to emerge on its own. We propose an architecture search framework that makes this alignment an explicit search variable: the assignment of data clusters to experts is optimised jointly with the per-expert architectures. We cast the joint problem as a cluster-aware likelihood maximisation, show that it coincides with the incomplete-data maximum likelihood of a latent-variable mixture, and solve it by a generalised Expectation-Maximisation procedure whose otherwise intractable expert-quality term is supplied by an adaptively refined surrogate. We prove that the iterates converge whenever the surrogate errors are summable, and that at every limit point no candidate the search produces improves the true objective. On a heterogeneous image-classification mixture the method recovers the underlying domain partition on 95% of clusters without ever observing domain labels, and on that benchmark and a four-domain time-series forecasting one alike it outperforms the MoE and NAS baselines that likewise use no label information.

Comment: 22 pages, 3 figures

arXiv abs page · PDF · same-day batch