Dropping a third of an AI’s chosen experts barely dented scores

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

TL;DR In nine AI models built from many small expert sub-networks, using only the top two thirds of each word’s chosen experts kept 98.8% of test performance on average. That one-setting change also sped up text generation, and fancier selection rules mattered only at deeper cuts.

The problem

Many open AI models are now a mixture of experts: lots of small sub-networks, called experts, plus a router that picks a handful for each word. Each pick costs computing time. Researchers have proposed rules for skipping some, word by word. But most were tested on older designs with few experts, and on multiple-choice quizzes where the model only ranks answers. Would the savings hold when a model must write a proof or program?

The idea

The team tested the simplest option first, then asked what cleverer rules add. Picture a cook whose sauce blends several stocks. Skip the smallest pours and scale up the rest, and diners barely notice until you skip most. The simple cut does that for every word alike: it keeps the experts the router weighted highest and rescales their weights. They swept this cut across nine models. On a subset, they pitted published per-word rules against it, at mild and deep cuts with matched average expert counts. Dropping a third of the experts, the authors explain, need not drop a third of their contribution.

What they showed

  1. Two thirds is enough. With that cut, the nine models kept 98.8% of their full scores on average, across knowledge, math, code and reasoning tests. Text generation also ran about 1.2 to 1.7 times faster on the models the team timed.
  2. Smarter rules add little at mild cuts. At that budget, the best published per-word rule scored within one percentage point of the simple cut on average test accuracy, and below it on some models.
  3. Deep cuts are where they pay. With far fewer experts left, the best rule won back up to about 3 points, mostly on tests where the model writes out full solutions, such as math and code.

Why it matters

Teams serving models like Qwen3 or GPT-OSS get an easy first step: the authors call it a one-number change to how many experts each word gets. Researchers designing pruning rules get a yardstick, since a new rule should beat this uniform cut at equal budget, especially on long answers. For you, it could mean faster answers from the same hardware, though price was never measured.

The catch

The small margins sit near the noise. Repeat runs of the same baseline differed by up to a point, about the size of the gaps at mild cuts. The deep-cut gains came with no speed tests, and per-word rules are harder to run fast, so their real payoff is unknown. The rule comparison covered only four models, and one rule was only partly rebuilt.

Who wrote it

First author Yuanteng Chen, corresponding authors Peisong Wang and Jian Cheng and seven colleagues, from the Institute of Automation, Chinese Academy of Sciences, Zhongguancun Academy and two others. Chen and Cheng co-authored “EAC-MoE” (2025), a compression method for such models.