Chapter 37
GRPO & Verifiable Rewards
Group-relative advantages, RLVR, and the DAPO, Dr. GRPO, and GSPO refinements.
Group Relative Policy Optimization (GRPO) is a critic-free reinforcement-learning recipe for post-training language models on tasks with checkable answers. For 2026 reasoning models, its appeal is practical: sample several answers to the same prompt, score them with a verifier, and increase the answers that beat their local group. DeepSeekMath used GRPO for mathematical reasoning, and DeepSeek-R1 made verifiable rewards central to reasoning post-training [shao2024] [deepseekai2025deepseekr1].
37.1 Groups and verifiable rewards
For each prompt , sample a group of answers from the old policy. A verifier returns scalar rewards . In RL with verifiable rewards, the verifier is not a learned preference model: it can be a unit test, an exact-match checker, or a parser that extracts a boxed arithmetic answer. The tiny checker below accepts only the exact integer sum for a prompt .
The group itself supplies the baseline. GRPO normalizes rewards inside the group:
If every answer in the group receives the same reward, all advantages are zero and the prompt contributes no policy signal. That is useful: a group with all wrong or all right answers cannot rank answers for this prompt.
Sampling more than one answer per prompt is what makes the baseline local. A hard prompt may have low absolute rewards for every model, and an easy prompt may have high absolute rewards, but GRPO does not compare those prompts directly. It asks which completion in this group was better than its siblings. The price is variance: if the group misses the rare correct answer, the verifier has nothing useful to rank. Production systems therefore care about prompt filtering, group size, and sampling temperature, not just the loss formula.
def arithmetic_reward(prompt, answer):
"""Return 1 when answer is the exact integer sum in a prompt (a, b)."""
try:
return float(int(str(answer).strip()) == int(prompt[0] + prompt[1]))
except ValueError:
return 0.0
def group_advantages(rewards, eps=1e-8):
"""A_i = (r_i - group_mean) / group_std for rewards shaped (B, G)."""
rewards = np.asarray(rewards, dtype=np.float64)
centered = rewards - np.mean(rewards, axis=1, keepdims=True)
std = np.std(rewards, axis=1, keepdims=True)
return np.where(std > eps, centered / (std + eps), 0.0)
37.2 PPO without a critic
GRPO keeps PPO’s clipped importance-ratio surrogate but removes the value network [schulman2017proximal]. For a sampled answer, define
The maximized surrogate is
A positive-advantage answer stops receiving extra credit once rises past the clip range; a negative-advantage answer stops being punished once its probability has fallen enough. That clipping makes several optimization epochs on the same sampled group less likely to move too far from the behavior policy.
GRPO also penalizes drift from a frozen reference policy. With samples from the current policy, the estimator from Section 7.8 estimates using :
The minimized loss is the negative clipped surrogate plus . Using rather than the raw log-ratio matters because individual samples are always nonnegative while the expectation is the same KL. A negative raw estimate can otherwise hide drift on a small batch. The KL term is still only an estimate, so it complements rather than replaces the PPO ratio clip.
def k3_kl_to_reference(logp, log_ref):
"""Schulman's k3 estimate of KL(policy || reference) for policy samples."""
log_ratio = log_ref - logp
return np.expm1(log_ratio) - log_ratio
def grpo_loss_and_grad(logits, old_logits, ref_logits, actions, advantages,
lengths, beta_kl=0.02, clip_eps=0.2,
normalize="sequence"):
"""PPO-clipped GRPO loss with a k3 KL penalty and no critic."""
logits = np.asarray(logits, dtype=np.float64)
logp = log_softmax(logits)
old_logp = log_softmax(old_logits)
ref_logp = log_softmax(ref_logits)
weights = normalization_weights(lengths, normalize)
grad_logp = np.zeros_like(logits)
loss = 0.0
for prompt in range(actions.shape[0]):
for sample in range(actions.shape[1]):
action = actions[prompt, sample]
advantage = advantages[prompt, sample]
weight = weights[prompt, sample]
ratio = np.exp(logp[prompt, action] - old_logp[prompt, action])
clipped = np.clip(ratio, 1.0 - clip_eps, 1.0 + clip_eps)
use_ratio = ((advantage >= 0.0 and ratio <= 1.0 + clip_eps)
or (advantage < 0.0 and ratio >= 1.0 - clip_eps))
objective = (ratio if use_ratio else clipped) * advantage
kl = k3_kl_to_reference(logp[prompt, action], ref_logp[prompt, action])
loss += weight * (-objective + beta_kl * kl)
if use_ratio:
grad_logp[prompt, action] -= weight * advantage * ratio
grad_logp[prompt, action] += weight * beta_kl * (
1.0 - np.exp(ref_logp[prompt, action] - logp[prompt, action]))
probs = np.exp(logp)
grad = grad_logp - probs * np.sum(grad_logp, axis=1, keepdims=True)
return float(loss), grad
37.3 Length and baselines
Language-model answers have different lengths, so an implementation must decide what an average means. Sequence-level normalization gives each sampled answer equal weight, often after summing or averaging token log-probabilities within that answer. Token-level normalization sums token losses across the batch and divides by the total number of generated tokens, so longer answers contribute more token terms. Neither convention is harmless: sequence-level weighting can hide length effects, while token-level weighting can make long completions dominate a batch.
The group standard deviation in (37.1) is also a design choice. It makes groups with small reward spread comparable to groups with large spread, but it changes the objective by rescaling each prompt. RLOO, the leave-one-out baseline, removes only the local mean: for sample , subtract the mean reward of the other answers. It keeps the baseline independent of while avoiding a learned critic.
RLOO and standard GRPO answer slightly different questions. RLOO says, "did this answer beat the other sampled answers?" in reward units. Standardized GRPO says, "how many within-group standard deviations better was it?" The second can stabilize mixed tasks whose reward scales differ, but it also changes the relative weight of prompts. That is why later variants spend so much attention on normalization.
def rloo_advantages(rewards):
"""Leave-one-out baseline: compare each reward with its peers' mean."""
rewards = np.asarray(rewards, dtype=np.float64)
group_size = rewards.shape[1]
if group_size < 2:
raise ValueError("RLOO needs at least two samples per prompt")
peer_sum = np.sum(rewards, axis=1, keepdims=True) - rewards
return rewards - peer_sum / (group_size - 1)
37.4 Recent refinements
DAPO keeps the GRPO family but adds a higher upper clip for positive advantages, dynamic sampling that filters uninformative groups, token-level policy-gradient loss, and overlong-answer shaping [yu2025dapo]. Dr. GRPO argues that standard GRPO’s length normalization can bias optimization toward longer incorrect answers, and removes length normalization; the same line of work also removes reward standard-deviation normalization when raw success-rate optimization is desired [liu2025understanding]. GSPO moves the importance ratio from token level to sequence level so the clipping decision follows the whole sampled answer [zheng2025group].
The common thread is not a new reward signal. It is better control over which sampled answers produce gradients and how those gradients scale with length, group composition, and probability ratio. For small examples, the original GRPO equations are enough; for large reasoning models, these normalization details become training behavior.
37.5 A tiny arithmetic run
The toy run has two prompts, each with three candidate answers. The group contains all candidates, the verifier gives reward one to the exact sum, and the policy starts uniform. Group-relative advantages push probability toward the correct answer for each prompt, while the KL penalty keeps some mass on the reference. The tests gradient-check the loss and assert that the trained policy puts most probability on the verifiable answer.
def toy_grpo_run(steps=80, lr=0.4):
"""Train two arithmetic prompts over three candidate answers each."""
prompts = [(1, 2), (2, 3)]
answers = np.array([["3", "4", "2"], ["5", "4", "6"]])
actions = np.tile(np.arange(3), (2, 1))
rewards = np.array([[arithmetic_reward(p, a) for a in row]
for p, row in zip(prompts, answers)])
advantages = group_advantages(rewards)
lengths = np.vectorize(len)(answers)
logits = np.zeros((2, 3), dtype=np.float64)
ref_logits = np.zeros_like(logits)
for _ in range(steps):
old_logits = logits.copy()
_, grad = grpo_loss_and_grad(logits, old_logits, ref_logits, actions,
advantages, lengths, beta_kl=0.03)
logits -= lr * grad
return np.exp(log_softmax(logits)), rewards
|
In practice
|
Modern RLVR pipelines rely on tasks where a checker is trusted more than a learned reward model: math, code, tool calls, or constrained formats [deepseekai2025deepseekr1]. Groups are sampled per prompt because a reward of one is informative only relative to competing answers from the same model. Reference KL is still used, but the main learning signal is often sparse success or failure. The main engineering questions are how to keep enough mixed-reward groups, how to handle very long answers, and how much KL to spend. |
37.6 Teach it
GRPO is PPO where the value baseline is replaced by classmates: compare each sampled answer with other answers to the same prompt. Analogy: grade a quiz by asking which of one student’s drafts passed the checker, not by predicting a global score. Board steps: sample answers; compute verifier rewards; center and scale within the group; apply the PPO clipped ratio and a reference KL penalty. Misconceptions: GRPO does not need a critic, but it still needs an old policy for ratios; a verifier is not automatically dense, because many groups can be all wrong; length normalization is part of the objective, not bookkeeping. Check: why does an all-wrong group have zero group-relative advantage?
37.7 Exercises
For rewards , compute the mean, standard deviation, and advantages. Explain why rewards produce no policy gradient.
For samples from , show that the expectation of (37.4) is .
Two completions have lengths and . Compute their weights under sequence-level and token-level normalization, and describe the training consequence.
Implement the arithmetic checker, group-relative advantages, clipped GRPO loss, and a tiny categorical training loop. Verify the gradient and show that correct answers gain probability.
References
-
[shao2024] Z. Shao et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. 2024. arXiv:2402.03300
-
[deepseekai2025deepseekr1] DeepSeek-AI et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. arXiv:2501.12948
-
[schulman2017proximal] J. Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347
-
[yu2025dapo] Q. Yu et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025. arXiv:2503.14476
-
[zheng2025group] C. Zheng et al. Group Sequence Policy Optimization. 2025. arXiv:2507.18071
-
[liu2025understanding] Z. Liu et al. Understanding R1-Zero-Like Training: A Critical Perspective. 2025. arXiv:2503.20783