PaperScope
LIVE · 2026-09-17 05:40 UTC

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

Eunju Shin, Jongbin Ryu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.18131 v1
Category
Submitted
2026-09-16

Abstract

In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q

Comment: Accepted by the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2026

arXiv abs page · PDF · same-day batch