RL should be the same shape as pretraining
RL for LLMs should be the same shape as pretraining. A training example is a sequence of tokens, a loss mask, and a per-token weight. A batch is packed sequences. The loss is weighted NLL:
where is the set of unmasked tokens in the batch and is the advantage.
Set and this is pretraining.
This means RL should run on the same infrastructure as pretraining: same packing, same batch API (tokens, masks, per-token weights), same metrics. gradnorm is just the gradient norm of a weighted NLL.
This is not just an engineering convenience. The math wants it to be this way. Three claims:
- Token-level RL is the right abstraction: the states and actions of the MDP are tokens-in-context, and the token-level gradient estimator is correct. Some recent work argues token-level importance ratios are wrong and you should use sequence-level ratios. We’ll derive what’s actually justified.
- The loss should be averaged over the total number of unmasked tokens in the batch, exactly like pretraining — not averaged per-sequence first.
- The right mental model is the classic RL replay buffer: a bag of transitions. A packed batch is just transitions laid out contiguously so they share prefix compute.
Let’s go straight to the math.
The token MDP
The goal is to find a policy that answers question with response , maximizing
The standard estimator is the policy gradient:
using the score-function identity to turn the gradient of an expectation into an expectation of a gradient.
So far this is sequence-level. But an LLM policy is autoregressive, so the sequence log-probability decomposes exactly:
This is the token MDP: the state is , the action is , and the transition is deterministic. No approximation has been made — the token-level estimator
is the policy gradient. Token-level is not a heuristic decomposition of a fundamentally sequence-level object; sequences are the trajectories of a token-level MDP.
Baselines
The raw estimator has high variance: easy problems always get reward 1 and dominate the update. We subtract a problem-dependent baseline :
This is unbiased because the score function has zero mean:
A Monte Carlo estimate of the expected reward is a valid baseline as long as it does not depend on the sample being weighted, giving the leave-one-out estimator. Defining (constant across for outcome rewards) and plugging into the token decomposition gives exactly the weighted NLL in equation (1).
Is token-level wrong off-policy?
On-policy, the derivation above settles it. The controversy is off-policy: sampling is expensive, so we take several gradient steps per batch, and the data was sampled from a stale policy . Recent work (GSPO) argues that the per-token ratio used by PPO/GRPO is not a valid importance weight, and that the correct weight is the sequence-level ratio.
The complaint is technically correct as far as it goes: the unbiased correction for using samples from is
the product of token ratios, not any individual one. But unbiased is not the right goal here — the variance of this product grows exponentially with sequence length, and clipping it discards entire sequences at a time.
The token-level estimator has its own justification, and it’s the classic one. The performance difference lemma writes the improvement of over as an expectation over states visited by ; the trust-region surrogate (CPI, TRPO, PPO) approximates that state distribution with ’s:
Here the per-token ratio is doing a precise job: it corrects the action distribution at each state the old policy visited. The only thing left uncorrected is the mismatch between state distributions and , and the surrogate’s error is bounded by the KL between the policies — which is exactly what clipping and KL regularization keep small.
So the choice is: unbiased with exponential variance (sequence-level), or biased by a state-distribution term that the trust region controls (token-level). Classic deep RL made this choice long ago — PPO on robotics tasks is the token-level surrogate over pairs — and it’s the right choice for LLMs for the same reason. Token-level ratios are not an LLM-era hack; they are the standard trust-region estimator applied to the token MDP.
Average over tokens, like pretraining
Equation (1) averages over all unmasked tokens in the batch. The original GRPO objective instead averages over tokens within each sequence, then averages over sequences. That gives every sequence equal weight, which means a token in a long sequence gets less weight than a token in a short one — an implicit length-dependent reweighting that produces the well-documented length biases (Dr. GRPO, DAPO).
The pretraining convention — sum the masked loss over the batch, divide by the total number of active tokens — gives every token equal weight. In the token MDP view this is the only natural choice: each token is a transition, and the estimator is a uniform average over transitions. Per-sequence averaging weights each transition by the inverse length of its episode, which is not an estimator of anything you wrote down. If you want length shaping, put it in the reward, not in the averaging convention.
This is also the convention your infrastructure already implements. If RL uses the pretraining loss reduction, packing changes nothing, gradient accumulation changes nothing, and metrics keep their interpretations.
A packed batch is a replay buffer
Classic RL got this shape right decades ago. A replay buffer doesn’t store episodes; it stores transitions , and a minibatch is a uniform sample of transitions. Nobody normalizes a DQN update by episode length.
The LLM analog: a transition is a token in context, an episode is a sequence, and a packed batch is a set of transitions laid out contiguously so that transitions from the same episode share prefix compute. Token-level averaging is uniform sampling from the buffer. The loss mask is just the mechanism for selecting which transitions in the buffer belong to the policy (sampled tokens) versus the environment (prompt tokens).
The buffer view also makes the off-policy story unsurprising: a replay buffer is off-policy by construction, which is precisely why the importance-weight discussion above exists, and why the trust-region surrogate — not the unbiased sequence ratio — is what every replay-based method actually optimizes.
RL is not a different kind of training that happens to reuse pretraining code. It is weighted NLL over a buffer of transitions, and pretraining infrastructure is already a transition-buffer trainer with all weights set to 1.
Next: MCTS is a replay buffer is a gradient estimator, where branching rollouts meet gradient estimator theory.