Why Routing Replay in MoE RL is Necessary
Reinforcement learning with Mixture-of-Experts (MoE) policies is inherently a latent-variable problem: at each timestep the policy samples a latent route and then an action , and that routing sequence affects all future states. If we discard the sampled routing path and try to train only on the marginal mixture over experts, the natural approximations give rise to a nested Monte Carlo estimator in the sense of Rainforth et al. (2018)1, which has provably poor statistical properties. The central claim of this note is that to avoid this nested-estimator pathology, one must perform routing replay: log the sampled routes and train on the resulting complete-data objective.
The remainder of the post unpacks this statement: we first formalize MoE RL as a latent-variable model (Sections 1–2), then characterize the nested estimator that arises when routes are marginalized and approximated (Section 4.1), and finally explain why routing replay restores a single-level, well-behaved estimator.
1. MoE policy as a latent-variable model
We consider an MoE policy in an RL setting with state , action , and reward at time . In addition to these observed variables, the policy includes a latent routing variable indicating which expert is selected at that timestep.
The router (gating network) defines a distribution over experts , and the expert policies define conditional action distributions .
The joint policy over is
The marginal action policy (what the environment actually sees) is the MoE mixture:
So is literally a latent variable in a mixture model: we do not feed it to the environment, but it controls which expert produced the action. In an actor–critic setup, a standard policy-gradient objective against this mixture is
where is an advantage estimate. Substituting the mixture form gives
Conceptually, this means:
- We are maximizing a marginal likelihood over actions in which the expert index is unobserved.
- Gradients must flow through a log-sum over experts, which is expensive and can be unstable when there are many experts or when routing is discrete.
This is exactly the same structure as maximum likelihood in a mixture model with unknown component assignments, and it is this marginal form that will give rise to a nested estimator when approximated.
1.1 Top- routing as a subset distribution
Many practical MoE architectures route not a single expert, but a subset of experts at each timestep. In that case the latent routing variable is a subset with , rather than a single index. Given router scores , a common abstraction is to view the induced distribution over orderings as a Plackett–Luce model:
where ranges over permutations of . Selecting the “top-” experts then corresponds to taking the first positions (possibly with additional deterministic or stochastic pruning).
In real systems, the router scores themselves are often subject to small amounts of implementation-level noise (for example, due to nondeterministic GPU kernels or reduction orderings), which effectively induces an implicit sampling distribution over experts even when the router is nominally deterministic.2 From the latent-variable perspective, this means that the subset is genuinely random under a distribution parameterized by the router, and routing replay is logging samples from this random subset distribution.
One concrete way to sample from this subset distribution is via the Gumbel–top- trick. Define nonnegative weights and draw i.i.d. Gumbel noise . Form perturbed scores and sort the experts by in descending order. It can be shown that the resulting random permutation is exactly distributed according to the Plackett–Luce model with parameters , so taking the first experts in this sorted order yields a size- subset whose distribution matches the Plackett–Luce–induced “top- without replacement” distribution.
2. Routing replay as complete-data training
We now relate this formulation to the practical routing replay technique used in RL. In routing replay, you store the router’s decision at data collection time:
More concretely:
- During data generation, is sampled from the behavior router .
- During training, you don’t infer from ; instead, you treat as observed data.
In latent-variable language, routing replay corresponds to choosing a degenerate inference model:
where is the stored routing choice in the replay buffer.
Then the inner expectation collapses:
because the entropy term is zero for a delta.
Plugging this into the ELBO-PG objective gives the routing replay objective:
This is exactly the complete-data objective you would write down for a mixture model where the component assignment (the expert index) is known.
3. No routing replay and the nested estimator
We can now address the central question: what goes wrong if we do not perform routing replay? From the latent-variable perspective, discarding the sampled routes and attempting to learn only from the marginal mixture
forces us to optimize
i.e., a marginal mixture likelihood with a discrete latent and a log-sum over experts.
This has several issues:
- Credit assignment to the router becomes poorly conditioned: every expert contributes inside the log-sum, even if only one expert actually produced the action.
- In realistic MoE architectures, the sum over experts is expensive or intractable, so it is natural to approximate it by Monte Carlo over routes; as we explain below, this introduces a nested Monte Carlo estimator with poor bias–variance behaviour (in the sense of Rainforth et al., 2018).
- The router can exploit the log-sum-exp structure in undesirable ways (for example, collapsing to a few experts or inducing brittle routing changes), because its gradients are driven by a “soft” mixture rather than the actual routing decisions used to act.
From a sequence-level RL perspective, the nesting can be seen explicitly. Consider an episodic setting with trajectories and a single return shared across all timesteps. If we keep the routes (routing replay), a complete-data policy-gradient estimator has the form
which is a single expectation over and therefore non-nested. If we discard and instead work with the marginal MoE policy, the exact marginal gradient can be written as
Approximating each inner sum over by Monte Carlo leads to an estimator of the form
where, for each trajectory and timestep ,
This matches the generic nested structure
with identified with trajectories , and with the inner objects defined by expectations over routes at each timestep. There is a single outer expectation over trajectories and, for each timestep , an inner expectation over route choices , all coupled through the shared return .
At a more abstract level, a nested Monte Carlo estimator arises whenever we wish to estimate a quantity of the form
and we approximate the inner expectation by Monte Carlo:
before applying the outer nonlinearity . The resulting estimator
is nested: it contains an inner Monte Carlo loop inside an outer one. From the latent-variable perspective, passing from the complete-data objective to its marginalized form and then approximating the inner expectation by Monte Carlo is precisely what introduces this nested structure. In this sense, marginalizing from to when we still have access to samples of is a weakening of the estimator: it trades a single-level objective on complete data for a nested objective on partial data.
In contrast, routing replay keeps the complete-data objective from Section 2:
which has several advantages:
- Clean credit assignment: the router is trained on the actual route that was taken when the action was executed and the return was observed.
- A well-behaved gradient estimator: there is no need to differentiate through a log-sum over experts; only the log-probabilities of a single expert and its route are involved.
- Conceptual alignment with standard latent-variable training: the setting corresponds to an “E-step already done” regime in which component assignments are observed.
In summary, routing replay is the MoE RL analogue of training on complete data instead of marginalizing over unknown latents. If one discards the router decisions and trains only on the marginal mixture, one implicitly tackles a more difficult latent-variable optimization problem whose natural approximations are nested estimators with poor statistical properties.
This perspective helps explain why, in practice, MoE RL setups that do not log and replay routes often behave poorly, whereas routing replay tends to stabilize training: it replaces a nested, ill-conditioned latent-variable objective with a tractable complete-data one.
4. Quantitative analysis of the nested estimator
We now quantify how much worse nested estimators are compared to non-nested estimators. Rainforth et al. (2018)1 show that, for fixed total budget , nested Monte Carlo estimators of the form above are generally biased and exhibit worse mean-squared-error scaling than non-nested estimators, because the nonlinearity amplifies inner Monte Carlo error.
It is useful to contrast this with standard (non-nested) Monte Carlo. With a single expectation and total budget samples, the mean-squared error (MSE) of the usual Monte Carlo estimator scales as
In the nested setting, even with an optimal choice of inner and outer sample sizes subject to , Rainforth et al. show that
Equivalently, to reach a target accuracy , a non-nested estimator requires , whereas a nested estimator requires . This gap quantifies how much “weakening” the estimator by introducing an inner Monte Carlo loop degrades its statistical efficiency.
In the MoE RL setting, there are two natural granularities for . At the per-decision level, can be taken as and as the expert index , with given by the composition with the logarithm and policy-gradient weighting. At the trajectory level, corresponds to an entire trajectory and to the sequence of routing decisions . Crucially, in an MoE policy each is sampled before the action and affects all subsequent states via the dynamics . The inner expectation is therefore over an entire routing sequence whose dimensionality and influence both grow with horizon , so naive nested Monte Carlo here is genuinely “nested over time” and becomes particularly pathological in long-horizon MoE RL.
For fixed , define
Then the per-sample gradients for the marginalized and routing-replay objectives can be written as
By the law of total variance, conditioning on ,
The second term is non-negative, so in the idealized regime where exact marginalization over experts is tractable, the marginal objective yields a lower-variance gradient estimator than the complete-data objective. In practice, however, the exact sum over experts is often prohibitively expensive in large MoE architectures.
When the sum over experts is approximated by Monte Carlo, the marginal objective induces a nested Monte Carlo estimator. Writing
one replaces by a finite-sample estimate
and uses inside an outer expectation over the replay distribution. This is exactly the nested structure studied by Rainforth et al. (2018): a nonlinear transformation of an inner Monte Carlo estimate inside an outer expectation.
The key implication is that, for a fixed total sample budget, such nested estimators exhibit worse bias and mean-squared-error scaling than non-nested estimators. Thus approximate marginalization over experts can be substantially less statistically efficient than either: (i) exact marginalization (when feasible), or (ii) complete-data training with routing replay, which remains non-nested.
From this viewpoint, the primary benefit of routing replay is that it avoids nested Monte Carlo altogether by converting the latent route into observed data and optimizing a single-level objective.
4.1 Case study: variance behavior for MoE RL
It is also useful to compare, at a qualitative level, how the variance of different MoE RL gradient estimators behaves as a function of the horizon and the choice of estimator. Consider an episodic setting with a single return per trajectory and define, for simplicity, the per-timestep complete-data contribution
so that the routing-replay gradient estimator is
-
Case 1: Complete-data (routing replay), single sample per trajectory.
Even without nesting, variance typically grows at least linearly with . Writingand assuming each has non-zero variance and covariances are not strongly negative, we obtain for some . This is the usual “variance grows with horizon” behavior of sequence-level REINFORCE.
-
Case 2: Exact marginalization over experts (intractable ideal).
If we could compute the sum over experts exactly, the marginal gradient estimator would have lower variance than the complete-data one by the law of total variance: integrating out removes conditional variability. However, this regime is not computationally accessible in realistic MoE architectures. -
Case 3: Single-sample plug-in approximation to the marginal.
In practice, approximating with a single-sample plug-in of the form (or its gradient) introduces additional noise: each per-timestep term becomes a nonlinear function of a single random route sample. Summing these noisy terms over and scaling by a shared again gives at least linear growth in variance with , but with larger per-timestep variance and bias compared to the complete-data case.
Taken together, these cases suggest the following picture: routing replay does not eliminate the usual horizon-related variance issues of RL, but it avoids the additional variance and bias introduced by approximating a high-dimensional marginal log-sum over experts with noisy single-sample plug-ins. The nested Monte Carlo viewpoint clarifies why the marginal training problem is statistically harder than the complete-data one, even when both ultimately rely on single-sample estimators in practice.
6. One-line summary
Routing replay turns the expert index from an unobserved latent variable in the mixture policy into an observed variable stored in replay, so the MoE RL training objective becomes a weighted complete-data (latent-variable) objective:
which is precisely the standard latent-variable formulation with known component assignments.