Workshop · Accepted ·

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade

NeurIPS 2026 Workshop on On-Device Intelligence · Workshop

Summary

When load balancing spreads routing too uniformly across experts, router scores become less useful for deciding which experts to prune. Domain-aware score allocation accounts for which capabilities each expert supports.

Research contribution

The paper characterizes expert pruning under over-dispersed routing and proposes Minimax Expert Score Allocation (MESA) to limit worst-case degradation across domains.

Scroll the figure sideways to read its labels, or open the full-size image below.

Lead figure

Three scatter plots compare WikiText-2 perplexity with GSM8K, MMLU and GPQA-Diamond accuracy across four pruning methods at r = 0.25 plus an unpruned baseline; correlations are positive or near zero.

Check the tasks you need to preserve: lower WikiText-2 perplexity does not consistently identify the more accurate pruning method in this comparison.

Perplexity does not reliably predict task accuracy across pruning configurations under over-dispersed routing. The comparison covers four pruning methods at r = 0.25 plus an unpruned baseline on GSM8K, MMLU and GPQA-Diamond.

Source: Berkcan Kapusuzoglu et al., When Load-Balancing Goes Too Far (2026), Figure 2. License: CC BY 4.0; rasterized without scientific edits.

Open full-size figure for labels and details.

All publications