Chapter 37

GRPO & Verifiable Rewards

Group-relative advantages, RLVR, and the DAPO, Dr. GRPO, and GSPO refinements.

Group Relative Policy Optimization (GRPO) is a critic-free reinforcement-learning recipe for post-training language models on tasks with checkable answers. For 2026 reasoning models, its appeal is practical: sample several answers to the same prompt, score them with a verifier, and increase the answers that beat their local group. DeepSeekMath used GRPO for mathematical reasoning, and DeepSeek-R1 made verifiable rewards central to reasoning post-training [shao2024] [deepseekai2025deepseekr1].

37.1 Groups and verifiable rewards

For each prompt xx, sample a group GG of answers y1,…,yGy_1, \ldots,y_G from the old policy. A verifier returns scalar rewards rir_i. In RL with verifiable rewards, the verifier is not a learned preference model: it can be a unit test, an exact-match checker, or a parser that extracts a boxed arithmetic answer. The tiny checker below accepts only the exact integer sum for a prompt (a,b)(a,b).

The group itself supplies the baseline. GRPO normalizes rewards inside the group:

Ai=ri−rˉsr+ϵ,rˉ=1G∑jrj.(37.1)A_i = \frac{r_i - \bar{r}}{s_r + \epsilon}, \quad \bar{r}=\frac1G\sum_j r_j .\tag{37.1}

If every answer in the group receives the same reward, all advantages are zero and the prompt contributes no policy signal. That is useful: a group with all wrong or all right answers cannot rank answers for this prompt.

Sampling more than one answer per prompt is what makes the baseline local. A hard prompt may have low absolute rewards for every model, and an easy prompt may have high absolute rewards, but GRPO does not compare those prompts directly. It asks which completion in this group was better than its siblings. The price is variance: if the group misses the rare correct answer, the verifier has nothing useful to rank. Production systems therefore care about prompt filtering, group size, and sampling temperature, not just the loss formula.

Listing 37.1 Verifiable reward and group-relative advantages
def arithmetic_reward(prompt, answer):
    """Return 1 when answer is the exact integer sum in a prompt (a, b)."""
    try:
        return float(int(str(answer).strip()) == int(prompt[0] + prompt[1]))
    except ValueError:
        return 0.0


def group_advantages(rewards, eps=1e-8):
    """A_i = (r_i - group_mean) / group_std for rewards shaped (B, G)."""
    rewards = np.asarray(rewards, dtype=np.float64)
    centered = rewards - np.mean(rewards, axis=1, keepdims=True)
    std = np.std(rewards, axis=1, keepdims=True)
    return np.where(std > eps, centered / (std + eps), 0.0)

37.2 PPO without a critic

GRPO keeps PPO’s clipped importance-ratio surrogate but removes the value network [schulman2017proximal]. For a sampled answer, define

ρi(θ)=exp⁡(log⁡πθ(yi∣x)−log⁡πold(yi∣x)).(37.2)\rho_i(\theta) = \exp\big(\log\pi_\theta(y_i|x)-\log\pi_{\mathrm{old}}(y_i|x)\big).\tag{37.2}

The maximized surrogate is

min⁡(ρiAi,clip⁡(ρi,1−ϵ,1+ϵ)Ai).(37.3)\min\big(\rho_i A_i, \operatorname{clip}(\rho_i,1-\epsilon,1+\epsilon)A_i\big).\tag{37.3}

A positive-advantage answer stops receiving extra credit once ρi\rho_i rises past the clip range; a negative-advantage answer stops being punished once its probability has fallen enough. That clipping makes several optimization epochs on the same sampled group less likely to move too far from the behavior policy.

GRPO also penalizes drift from a frozen reference policy. With samples from the current policy, the k3k_3 estimator from Section 7.8 estimates DKL(πθ∥π0)\KL(\pi_\theta\Vert\pi_0) using ui=π0(yi∣x)/πθ(yi∣x)u_i=\pi_0(y_i|x)/\pi_\theta(y_i|x):

k3(i)=(ui−1)−log⁡ui.(37.4)k_3(i) = (u_i - 1) - \log u_i .\tag{37.4}

The minimized loss is the negative clipped surrogate plus βKLk3\beta_{\mathrm{KL}} k_3. Using k3k_3 rather than the raw log-ratio matters because individual samples are always nonnegative while the expectation is the same KL. A negative raw estimate can otherwise hide drift on a small batch. The KL term is still only an estimate, so it complements rather than replaces the PPO ratio clip.

Listing 37.2 The clipped GRPO loss with k3 KL penalty
def k3_kl_to_reference(logp, log_ref):
    """Schulman's k3 estimate of KL(policy || reference) for policy samples."""
    log_ratio = log_ref - logp
    return np.expm1(log_ratio) - log_ratio


def grpo_loss_and_grad(logits, old_logits, ref_logits, actions, advantages,
                       lengths, beta_kl=0.02, clip_eps=0.2,
                       normalize="sequence"):
    """PPO-clipped GRPO loss with a k3 KL penalty and no critic."""
    logits = np.asarray(logits, dtype=np.float64)
    logp = log_softmax(logits)
    old_logp = log_softmax(old_logits)
    ref_logp = log_softmax(ref_logits)
    weights = normalization_weights(lengths, normalize)
    grad_logp = np.zeros_like(logits)
    loss = 0.0
    for prompt in range(actions.shape[0]):
        for sample in range(actions.shape[1]):
            action = actions[prompt, sample]
            advantage = advantages[prompt, sample]
            weight = weights[prompt, sample]
            ratio = np.exp(logp[prompt, action] - old_logp[prompt, action])
            clipped = np.clip(ratio, 1.0 - clip_eps, 1.0 + clip_eps)
            use_ratio = ((advantage >= 0.0 and ratio <= 1.0 + clip_eps)
                         or (advantage < 0.0 and ratio >= 1.0 - clip_eps))
            objective = (ratio if use_ratio else clipped) * advantage
            kl = k3_kl_to_reference(logp[prompt, action], ref_logp[prompt, action])
            loss += weight * (-objective + beta_kl * kl)
            if use_ratio:
                grad_logp[prompt, action] -= weight * advantage * ratio
            grad_logp[prompt, action] += weight * beta_kl * (
                1.0 - np.exp(ref_logp[prompt, action] - logp[prompt, action]))
    probs = np.exp(logp)
    grad = grad_logp - probs * np.sum(grad_logp, axis=1, keepdims=True)
    return float(loss), grad

37.3 Length and baselines

Language-model answers have different lengths, so an implementation must decide what an average means. Sequence-level normalization gives each sampled answer equal weight, often after summing or averaging token log-probabilities within that answer. Token-level normalization sums token losses across the batch and divides by the total number of generated tokens, so longer answers contribute more token terms. Neither convention is harmless: sequence-level weighting can hide length effects, while token-level weighting can make long completions dominate a batch.

The group standard deviation in (37.1) is also a design choice. It makes groups with small reward spread comparable to groups with large spread, but it changes the objective by rescaling each prompt. RLOO, the leave-one-out baseline, removes only the local mean: for sample ii, subtract the mean reward of the other G−1G-1 answers. It keeps the baseline independent of rir_i while avoiding a learned critic.

RLOO and standard GRPO answer slightly different questions. RLOO says, "did this answer beat the other sampled answers?" in reward units. Standardized GRPO says, "how many within-group standard deviations better was it?" The second can stabilize mixed tasks whose reward scales differ, but it also changes the relative weight of prompts. That is why later variants spend so much attention on normalization.

Listing 37.3 RLOO advantages
def rloo_advantages(rewards):
    """Leave-one-out baseline: compare each reward with its peers' mean."""
    rewards = np.asarray(rewards, dtype=np.float64)
    group_size = rewards.shape[1]
    if group_size < 2:
        raise ValueError("RLOO needs at least two samples per prompt")
    peer_sum = np.sum(rewards, axis=1, keepdims=True) - rewards
    return rewards - peer_sum / (group_size - 1)

37.4 Recent refinements

DAPO keeps the GRPO family but adds a higher upper clip for positive advantages, dynamic sampling that filters uninformative groups, token-level policy-gradient loss, and overlong-answer shaping [yu2025dapo]. Dr. GRPO argues that standard GRPO’s length normalization can bias optimization toward longer incorrect answers, and removes length normalization; the same line of work also removes reward standard-deviation normalization when raw success-rate optimization is desired [liu2025understanding]. GSPO moves the importance ratio from token level to sequence level so the clipping decision follows the whole sampled answer [zheng2025group].

The common thread is not a new reward signal. It is better control over which sampled answers produce gradients and how those gradients scale with length, group composition, and probability ratio. For small examples, the original GRPO equations are enough; for large reasoning models, these normalization details become training behavior.

37.5 A tiny arithmetic run

The toy run has two prompts, each with three candidate answers. The group contains all candidates, the verifier gives reward one to the exact sum, and the policy starts uniform. Group-relative advantages push probability toward the correct answer for each prompt, while the KL penalty keeps some mass on the reference. The tests gradient-check the loss and assert that the trained policy puts most probability on the verifiable answer.

Listing 37.4 Tiny categorical GRPO on arithmetic answers
def toy_grpo_run(steps=80, lr=0.4):
    """Train two arithmetic prompts over three candidate answers each."""
    prompts = [(1, 2), (2, 3)]
    answers = np.array([["3", "4", "2"], ["5", "4", "6"]])
    actions = np.tile(np.arange(3), (2, 1))
    rewards = np.array([[arithmetic_reward(p, a) for a in row]
                        for p, row in zip(prompts, answers)])
    advantages = group_advantages(rewards)
    lengths = np.vectorize(len)(answers)
    logits = np.zeros((2, 3), dtype=np.float64)
    ref_logits = np.zeros_like(logits)
    for _ in range(steps):
        old_logits = logits.copy()
        _, grad = grpo_loss_and_grad(logits, old_logits, ref_logits, actions,
                                     advantages, lengths, beta_kl=0.03)
        logits -= lr * grad
    return np.exp(log_softmax(logits)), rewards
In practice

Modern RLVR pipelines rely on tasks where a checker is trusted more than a learned reward model: math, code, tool calls, or constrained formats [deepseekai2025deepseekr1]. Groups are sampled per prompt because a reward of one is informative only relative to competing answers from the same model. Reference KL is still used, but the main learning signal is often sparse success or failure. The main engineering questions are how to keep enough mixed-reward groups, how to handle very long answers, and how much KL to spend.

Key equations
Ai=(ri−rˉ)/(sr+ϵ)A_i = (r_i - \bar{r})/(s_r+\epsilon)
ρi=exp⁡(log⁡πθ(yi)−log⁡πold(yi))\rho_i = \exp(\log\pi_\theta(y_i)-\log\pi_{\mathrm{old}}(y_i))
Li=−min⁡(ρiAi,clip⁡(ρi,1−ϵ,1+ϵ)Ai)L_i = -\min(\rho_i A_i,\operatorname{clip}(\rho_i,1-\epsilon,1+\epsilon)A_i)
k3=(u−1)−log⁡u,u=π0(yi)/πθ(yi)k_3 = (u-1)-\log u,\quad u=\pi_0(y_i)/\pi_\theta(y_i)
AiRLOO=ri−1G−1∑j≠irjA_i^{\mathrm{RLOO}} = r_i - \frac{1}{G-1}\sum_{j\ne i} r_j

37.6 Teach it

GRPO is PPO where the value baseline is replaced by classmates: compare each sampled answer with other answers to the same prompt. Analogy: grade a quiz by asking which of one student’s drafts passed the checker, not by predicting a global score. Board steps: sample GG answers; compute verifier rewards; center and scale within the group; apply the PPO clipped ratio and a reference KL penalty. Misconceptions: GRPO does not need a critic, but it still needs an old policy for ratios; a verifier is not automatically dense, because many groups can be all wrong; length normalization is part of the objective, not bookkeeping. Check: why does an all-wrong group have zero group-relative advantage?

37.7 Exercises

Exercise 37.1 ★ Compute group-relative advantages

For rewards (1,0,0)(1,0,0), compute the mean, standard deviation, and advantages. Explain why rewards (0,0,0)(0,0,0) produce no policy gradient.

Exercise 37.2 ★★ Show what k3 estimates

For samples from πθ\pi_\theta, show that the expectation of (37.4) is DKL(πθ∥π0)\KL(\pi_\theta\Vert\pi_0).

Exercise 37.3 ★★ Compare length normalizations

Two completions have lengths 22 and 66. Compute their weights under sequence-level and token-level normalization, and describe the training consequence.

Exercise 37.4 ★★★ Implement the toy RLVR run

Implement the arithmetic checker, group-relative advantages, clipped GRPO loss, and a tiny categorical training loop. Verify the gradient and show that correct answers gain probability.

References

  • [shao2024] Z. Shao et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. 2024. arXiv:2402.03300

  • [deepseekai2025deepseekr1] DeepSeek-AI et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. arXiv:2501.12948

  • [schulman2017proximal] J. Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347

  • [yu2025dapo] Q. Yu et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025. arXiv:2503.14476

  • [zheng2025group] C. Zheng et al. Group Sequence Policy Optimization. 2025. arXiv:2507.18071

  • [liu2025understanding] Z. Liu et al. Understanding R1-Zero-Like Training: A Critical Perspective. 2025. arXiv:2503.20783