Adding a little noise to MoE routers helps the model find solutions to hard problems

We take a frozen mixture-of-experts reasoning model, and add Gaussian noise to the weights of its router.

Surprisingly, this simple, almost-free intervention results in a modest but consistent increase in performance for the problems on the edge of a model’s capability, precisely the ones that are most important for test-time scaling, RL post-training, and continual learning.

Figure 1. Hard-set Pass@k on GPT-OSS-120B for the stock model vs. the noisy router (pn = 0.02). Each curve estimated from 64 samples, hard-set defined by 64 held-out stock samples.

This blog post explains how we arrived at this striking result: our perspective on search, the increasing popularity of MoE models and their connections to ensembles in classical ML, and the independent reasons from neuroscience for this experiment.

MoEs give us an unusually explicit search surface

A sparse mixture-of-experts model contains many feed-forward "experts," but only a small subset processes any given token.

A learned router matrix makes that choice: For a token representation h, the router scores the available experts and selects the top few. Those experts transform the representation; the rest sit unused for that routing decision.

Figure 2. Hidden state → router → large expert pool → top-k active experts. The router scores all 32 experts and recruits the top four; the unused experts are ghosted.

That makes MoE routing a natural place to ask:

Could some useful trajectories live behind alternative routing decisions the stock router almost never makes?

This makes MoEs interesting from a search perspective: the frozen model already contains multiple possible computational paths. The experts already exist, and their weights have already been learned. What changes from token to token is which subset gets recruited, a decision controlled by the router matrices.

There is an analogy here to model ensembles. Ensembles can outperform their individual members when the members make uncorrelated errors. Although an MoE is not an ensemble of complete models, there is a related structural fact: it contains a collection of alternative computations, and the router determines which of them participate.

That suggests a cheap source of diversity. Rather than running an entirely different model, perhaps we can induce different computation by changing routing inside the same model.

There was another clue we found hard to ignore: modern MoEs are becoming increasingly sparse.

Mixtral, released in late 2023, activates 2 of 8 experts (25%); DeepSeek-V3, released just a year later, activates 8 of 256 routed experts (3.1%). And most recently, the frontier Kimi K3, released in mid-2026, activates only 16 of 896 (1.8%).

Figure 2b. The sparsity of an earlier MoE model (Mixtral) against later ones (DeepSeek-V3, Kimi K3), with every expert pool drawn to the same scale, one dot per expert. Filled dots are the experts one token activates; each new token activates a different few.

This gave us a question we thought worth testing:

As expert pools grow while activation stays sparse, does the trained router leave more useful computation unexplored?

Before getting to the intervention, there was one more reason we were interested in that question.

Neuroscience motivation: How variability changes which circuits participate

Neuroscience gave us another reason to investigate routing. A brain circuit can produce different responses without changing its anatomical connections. The inputs arriving at a given moment, and how the circuit responds to them, influence which groups of neurons participate. For an MoE, this suggested a question: could changing how experts are recruited help a frozen model explore useful alternative computations?

To make the connection concrete, consider a small local population of neurons receiving inputs from several sources. Some inputs excite the population; others recruit inhibition that restrains it. Whether the population becomes strongly active depends on their combined influence. In this simplified picture, recruitment means that the population’s activity becomes strong enough to influence subsequent processing. A microcolumn is one possible anatomical picture to have in mind, but we use the population as an illustrative computational unit, without assuming that cortical microcolumns function as discrete experts.

There are two ways its recruitment could change. First, the same combination of inputs could become more or less effective because the population becomes more responsive or less inhibited. Second, the balance among inputs could change: one source might exert more influence than another, changing which combinations of inputs recruit the population. The first changes how readily the population responds; the second changes what it responds to. Both can alter its participation without requiring new anatomical connections.

Biological experiments illustrate parts of this picture. Synaptic transmission itself is variable: neurotransmitter release is probabilistic, so repeated activity can deliver different inputs through the same connection (Jensen et al., 2021). Those fluctuations enter a system whose responsiveness can also change. Chance et al. (2002) supplied cortical neurons with simulated, fluctuating excitation and inhibition. Increasing both together made firing rates rise less steeply with additional input. The background activity changed how the neurons responded to incoming signals.

Chemical signals can adjust responsiveness through particular circuit interactions. In mouse visual cortex, locomotion activated VIP inhibitory neurons largely through cholinergic input. These neurons suppressed another inhibitory population, allowing nearby excitatory neurons to respond more strongly to visual stimuli (Fu et al., 2014). Inhibiting an inhibitor made those neurons more responsive. Other experiments show selective recruitment: orbitofrontal and thalamic inputs preferentially activated different subnetworks within visual cortex while recruiting inhibition of the competing subnetwork (Liu et al., 2024). And changing the balance can matter for behavior: manipulating particular superior colliculus cell types shifted choice-related activity in frontal cortex and biased whether mice subsequently licked left or right (Thomas et al., 2023).

An MoE router makes a related operation explicit. It combines features of the current token representation into a score for each expert, then selects the highest-scoring experts. Perturbing the router changes how those features contribute to the scores. For a particular input, an expert may move into or out of the selected group. Across inputs, different combinations of features may now favor that expert. The expert itself remains unchanged.

This matters even when rollouts already vary because of token sampling. Different token histories produce different internal representations. Changing the router changes how those differences translate into expert selections, potentially amplifying some distinctions or suppressing others. Holding a perturbation fixed throughout an attempt applies a common altered rule to the changing representations. Applying it only at selected moments limits how often that rule replaces ordinary routing.

The biological evidence motivates this focus on responsiveness and recruitment. Whether the resulting variation improves search, and whether persistent or intermittent interventions work best, are questions for the model experiments. So we tried a simple intervention: adding Gaussian noise to MoE router weights while keeping the experts fixed.

So we tried the simplest version we could think of: adding Gaussian noise to the router matrix of MoE models.

The right dose of router noise improves performance on hard problems

In our experiments, we observe that the right dose of noise provides a noticeable improvement in hard-set Pass@32, defined as the set of problems with a success rate of ≤ 10% estimated from 64 held-out baseline samples:

Figure 3. Hard-set Pass@k (left) and Pass@32 by configuration (right) on GPT-OSS-120B. Bars show mean Pass@32 with ±1 std; 64 samples per condition.

This is an initial indication of more efficient search behavior for hard problems.

Although the improvement is modest, the interesting part is that an almost embarrassingly simple, training-free change to a tiny part of inference, at no additional inference FLOPs, reliably moves performance at all.

To measure the effect of router noise for easier problems, we also compute Pass@1 for the medium difficulty (between 10% and 80% success rate) and easy difficulty (>80% success rate) problem sets:

Figure 4. Pass@1 on the medium and easy problem sets for the noisy router and stock (mean ±1 std).

Surprisingly, the benefit to the hard problems does not come at a cost for the medium or easy problems.

A small amount of noise is beneficial; large amounts are harmful

However, the story is not as simple as "more noise is always better". We find that a key factor is the intervention probability: the amount of routing decisions changed from the baseline model. We plot Pass@32 at different amounts of intervention and compare to the stock router on the hard set:

Figure 5. GPT-OSS-120B intervention-frequency sweep on the hard set, 64 samples per condition.

Even very sparse interventions, around 1–2%, are enough to provide a noticeable improvement in hard-set Pass@32. At larger intervention rates, benefits are more inconsistent, and more often than not perform worse than baseline.

We see the same pattern on the medium difficulty (between 10% and 80% success rate) and easy difficulty (>80% success rate) problem sets, here shown with Pass@1:

Figure 6. GPT-OSS-120B intervention-frequency sweep, 64 samples per condition, medium and easy sets.

Together, these curves suggest a narrow window of useful router noise rather than “more exploration is always better.” In the sections that follow, we spell out the exact intervention method and how intervention probability is applied during inference.

Gaussian router perturbation

The dose–response curves above compare configurations indexed by pn. Here we describe the change to the router when a decision is perturbed.

At each MoE layer, the model has a learned router matrix W. Ignoring a few architectural details, we can write its output logits as

z = W h,

where h is the current hidden state and W is the learned router matrix. The router ranks the experts using z, then activates the top k.

Our intervention is:

z′ = z + ΔW h,

where ΔW is a matrix of Gaussian noise. Importantly, for any given trajectory, ΔW is sampled once and held fixed throughout the trajectory.

Figure 7. h → W + ΔW → expert logits → top-k experts, for the stock router and the perturbed router. The perturbation nudges every logit a little; here it lifts expert 8 above the top-2 cut and drops expert 5 below it.

That is nearly the whole method. The perturbation is applied only to the router. We do not modify the experts themselves. We do not touch attention, embeddings, or any of the other learned parameters, and no additional training is required.

There is one property of this intervention that becomes important later: We are not directly saying "replace expert 3 with a random expert." We perturb the function that scores experts, then hold that function fixed for an entire trajectory.

For one hidden state, ΔW h might raise one expert and lower another. For a different hidden state, the exact same ΔW can have a different effect. The change in routing therefore depends on what the model is currently processing.

We can think of each sampled ΔW as defining a nearby routing policy around the one the model learned.

But our first attempts raised an immediate problem: If changing the router sometimes exposes useful alternatives, should we change it all the time?

Intervention probability and gating

We have already seen that performance depends sharply on how often routing is perturbed. Why not intervene on every decision, or never?

Randomness and exploration are crucial for search, and the observations we outline above motivate a new axis of exploration via expert selection. However, naive, repeated intervention can be destructive.

A random perturbation may move the model toward a useful alternative computation, while perturbing the immediately following states can disrupt that computation before the model's learned dynamics can propagate, refine, or correct it. Because autoregressive computation conditions on the consequences of previous decisions, sparse interventions may provide diversity while allowing most of the trajectory to evolve under the model's learned computation.

We therefore study the effect of intervention probability: For every token at every MoE layer, we sample a Bernoulli gate with intervention probability pn. If the gate is off, routing is completely normal. If it is on, that routing decision uses the perturbed router.

Animation. The intervention probability moves from pn = 0% → 4% → 100%, and back to 4%, over a token × layer grid of routing decisions. A cell is green while its gate sends it to the perturbed router, so perturbed cells grow denser as pn rises.

This gave us a continuum between two extremes. At one end, nothing changes. We get no new routing behavior. At the other end, we continuously interfere with a router the model spent a great deal of compute learning.

Neither extreme seems right, which is why the sweeps in the preceding section vary the intervention probability pn between them.

What kinds of problems benefit the most?

We observed that the benefit is concentrated on hard problems. However, do all hard problems benefit equally?

Among the hard set, there are 33 problems where the noisy router model achieved a higher success rate compared to the stock model, but there are also 24 problems where stock outperformed the noisy router:

Number of hard-set problems where noisy router noisy router at pn = 0.02 (green, left) versus stock (red, right). Problems with Δ = 0 are omitted.

To answer this question more concretely, we compute the per-problem difference in success rate of the best-performing configuration, pn = 0.02 which intervenes with a noisy router in 2% of all routing decisions, against the stock model, estimated from 64 samples each. A green bar means that the noisy model did better than stock for that problem in our sample, and vice-versa for a red bar:

Figure 8. Per-problem difference in log-success rate between pn = 0.02 and the stock model (64 samples each). Green: noisy router better; red: stock better. Problems sorted by increasing difficulty (hardest at left), using 64 held-out stock samples.

Problem IDs are sorted in increasing difficulty from left to right, but estimated from 64 held-out stock samples to avoid in-sample bias and regression-to-mean effects. Additionally, since an improvement from 10% to 20% success rate is much more impactful than an improvement from 70% to 80%, we report the difference in log-success rate instead of raw success rate in percentage points.

Qualitatively, we see green dominating red in the hardest third of the benchmark (left end), explaining the noisy router model's advantage for the hard problems.

However, the advantage seems to be problem-specific, as there are some hard problems that the stock model does better at (red bars). We also note that there is a large amount of sampling noise as per-problem success rates may fluctuate widely even between two 64-sample sets of the same configuration (e.g. stock vs stock). Combined with the small effect of the method, it's unclear if there is a consistent advantage over the stock model.

To make this advantage precise, we plot the cumulative difference in log-success rates. This is done by summing the “area” of the green and red bars as each new problem is added: problems where pn = 0.02 is better adds to the cumulative value; problems where stock is better subtracts from it. If the noisy model truly is better than stock, we expect this cumulative sum to climb to a positive value; under the null hypothesis, the cumulative sum should remain close to zero.

Since exact problem ordering is arbitrary within a given broad difficulty band, we smooth the curve by taking the average over 20 problems at any given location on the x-axis. This results in the plot below:

Figure 9. Cumulative difference in log-success rate between pn = 0.02 and the stock model, summing per-problem differences in difficulty order (hardest first). The line is a moving average of that cumulative sum over a ±10-problem window. Problems sorted using 64 held-out stock samples.

We observe that the cumulative sum shows an overall positive trend up to around the top 120 hardest problems. But perhaps more surprisingly, the cumulative sum tapers off but does not go back down, showing that the noisy router model does not harm problems that the stock model already finds easy.

Fixed perturbation per trajectory or resampled perturbation?

Equally interesting is what we tried that did not work. Here we ablate a design choice made previously: the choice of sampling ΔW once and fixing it for a full trajectory.

First, we experiment with resampling a fresh ΔW at every routing decision at pn = 0.02 and compare it with the frozen ΔW system. Conceptually, this distinguishes between stochastic sampling at the level of individual routing decisions, vs sampling over a distribution of coherent routing policies.

This results in the following plot:

Figure 10. Cumulative difference in log-success rate between per-decision resampled ΔW and the stock model, summing per-problem differences in difficulty order (hardest first). The line is a moving average of that cumulative sum over a ±10-problem window.

Unlike the previous results we saw, adding randomness at the level of individual routing decisions achieves no noticeable advantage over the stock model on the hardest 120 problems, with the cumulative log advantage fluctuating around zero throughout most of this regime. This tells us something about the structure of hard reasoning problems: exploring alternative routing policies is genuinely useful, but maintaining coherent policies is an important factor over simply introducing unstructured stochasticity.

However, there is a clean positive trend starting from around the 110th problem all the way up to the remaining medium and easy problems. This is an interesting observation: structured sampling at the routing policy level seems to help on the hardest problems, while unstructured sampling at the level of individual routing decisions helps more on medium-hard problems.

This is made more clear by plotting the cumulative difference plot while treating the fixed-ΔW configuration as the baseline:

Figure 11. Cumulative difference in log-success rate between per-decision resampled ΔW and the fixed-ΔW configuration (both at pn = 0.02), summing per-problem differences in difficulty order (hardest first). The line is a moving average of that cumulative sum over a ±10-problem window (64 samples per condition).

We observe that the cumulative advantage clearly trends negatively for the first ~120 problems, indicating that the fixed-ΔW configuration maintains a clear advantage for these 120 hardest problems, but the trend reverses thereafter.

This indicates that both types of randomness play distinct roles, and a system that utilizes both can outperform the results we have seen today.

What we think this experiment says

The most surprising result here is how little we had to change.

We did not retrain the model, modify its experts, or add another search procedure around it. We only perturbed the router, and only exposed roughly 1–2% of all routing decisions to that perturbation. Yet that was enough to materially improve search performance on the hardest problems—the ones sitting closest to the edge of the model's capabilities.

The intervention is also unusually simple. Adding Gaussian noise to the router requires no additional inference FLOPs, no learned auxiliary model, and no modification to attention or the experts themselves. The model already contains the alternative computations; the noise only changes which ones get recruited.

More broadly, this experiment is a good example of the kind of synthesis we want to pursue at Voaige. We started from the view that search is central to intelligence, and independently, advancements in modern LLMs provided an unusually explicit internal search surface in the form of expert routing. Neuroscience gave us a third reason: that a closely related phenomenon occurs in the brain. All these factors converged to the same experiment, and that's why we pursued it.

Future work. The result is encouraging, but still preliminary. We have only tested one model, GPT-OSS-120B, and only on mathematical reasoning problems. The gains are also small relative to the amount of sampling noise in these evaluations. At the per-problem level, there are clearly cases where router noise appears to hurt, and with the evidence we have today we cannot reliably predict which problems those will be.

That leaves several important questions open. Does this effect generalize to other MoE architectures, domains, and forms of reasoning? Can we predict when routing perturbations will help rather than hurt? Can we design better perturbations than isotropic Gaussian noise, or adapt the amount and type of intervention to the problem and state of the trajectory? And can the effect be amplified enough to become a more substantial improvement rather than a small shift in search efficiency?

For now, the main takeaway is initial evidence toward a hypothesis: the router appears to matter as a search surface in its own right. A tiny change to how a frozen model recruits its experts can expose useful reasoning trajectories that ordinary sampling reaches less often. That is a small result, but it suggests there may be substantially more inference-time computation to explore inside the model than its default routing policy reveals.

References

  • Jensen, T. P., et al. (2021). Release probability increases towards distal dendrites boosting high-frequency signal transfer in the rodent hippocampus. eLife, 10, e62588. Read paper
  • Chance, F. S., Abbott, L. F., & Reyes, A. D. (2002). Gain modulation from background synaptic input. Neuron, 35(4), 773–782. Read paper
  • Fu, Y., et al. (2014). A cortical circuit for gain control by behavioral state. Cell, 156(6), 1139–1152. Read paper
  • Liu, Y., et al. (2024). Organization of corticocortical and thalamocortical top-down inputs in the primary visual cortex. Nature Communications, 15, 4495. Read paper
  • Thomas, A., et al. (2023). Superior colliculus bidirectionally modulates choice activity in frontal cortex. Nature Communications, 14, 7358. Read paper
Continue exploring

Research on cognition, search, and test-time compute.