PaperScope
LIVE · 2026-09-29 05:40 UTC

Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

Jinfan He, Yunzhuo Liu, Kai Zhang, Weidong Han, Key, Rayying

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32398 v1
Category
Submitted
2026-09-26

Abstract

The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.

Comment: COLM 2026 accept

arXiv abs page · PDF · same-day batch