Chapter 35
Reward Models, PPO & RLHF
Bradley-Terry reward models, the clipped PPO objective, and KL-regularized RLHF.
RLHF turns human or AI preferences into a reward, then uses reinforcement learning to move a language-model policy toward high-reward responses without drifting too far from a reference model. The pipeline is three pieces: learn a reward model from comparisons, define a KL-regularized objective, and optimize the policy with a stable policy-gradient method. PPO is the standard small step in that last piece.
35.1 Reward models from pairwise preferences
A preference dataset contains pairs: for the same prompt, response was preferred to response . A Bradley-Terry model says the probability of that preference is a sigmoid of reward difference [bradley1952]:
For one pair with , the negative log-likelihood is
Differentiating gives . For a linear reward , the gradient is . The code builds synthetic pairs from a hidden linear reward, gradient-checks this loss, and trains a small reward model to score winners above losers.
Only differences matter. Adding the same constant to both rewards leaves unchanged, so a reward model’s absolute zero is arbitrary. That is fine for policy optimization, which needs to rank sampled responses, but it means reward values should not be read like calibrated human scores. The synthetic data keeps one prompt implicit and uses feature vectors for responses; real systems condition the reward model on both prompt and response.
def bradley_terry_loss_and_grad(weights, winners, losers):
"""Loss -log sigmoid(r_w - r_l) for a linear reward model."""
weights = np.asarray(weights, dtype=np.float64)
features = np.asarray(winners) - np.asarray(losers)
margins = features @ weights
loss = np.mean(np.logaddexp(0.0, -margins))
sigmoid = 1.0 / (1.0 + np.exp(-margins))
grad = ((sigmoid - 1.0)[:, None] * features).mean(axis=0)
return float(loss), grad
def train_reward_model(winners, losers, steps=300, lr=0.5):
weights = np.zeros(winners.shape[1], dtype=np.float64)
for _ in range(steps):
_, grad = bradley_terry_loss_and_grad(weights, winners, losers)
weights -= lr * grad
return weights
35.2 KL-regularized RLHF
Once a reward model exists, the policy objective for a prompt is
The reward term pulls the policy toward preferred responses. The KL term keeps it close to a reference policy, usually the SFT model from Chapter 33, so reward-model mistakes do not dominate. The coefficient is the exchange rate between reward and drift.
The KL direction matters. Because responses are sampled from the current policy, the penalty is naturally estimated under , giving . This is the same reverse direction discussed in Section 7.4.1: it punishes probability mass the new policy puts where the reference was unlikely. A large keeps style and coverage close to the reference; a small lets the reward model steer more aggressively.
The full response-level KL sums over impossible-to-enumerate outputs. In practice, algorithms estimate a per-token KL on sampled responses; Section 7.8 derives common estimators and their variance behavior. The tiny categorical code computes the exact KL so the objective can be tested without sampling noise.
This objective is not yet an algorithm. It says which expectation to improve, but a language model cannot enumerate every response and take an exact gradient. PPO supplies the sampled, local update rule: collect completions from the current policy, freeze that policy as , estimate advantages, then take a few cautious gradient steps on the same batch.
def categorical_kl(policy_probs, reference_probs):
"""KL(pi || pi_ref) for categorical policies."""
policy_probs = np.asarray(policy_probs, dtype=np.float64)
reference_probs = np.asarray(reference_probs, dtype=np.float64)
log_ratio = np.log(policy_probs) - np.log(reference_probs)
return float(np.sum(policy_probs * log_ratio))
def rlhf_objective(policy_probs, rewards, reference_probs, beta):
"""E_pi[r] - beta KL(pi || pi_ref)."""
reward_term = float(np.asarray(policy_probs) @ np.asarray(rewards))
return reward_term - beta * categorical_kl(policy_probs, reference_probs)
35.3 PPO’s clipped surrogate
A policy-gradient update uses samples from an old policy but evaluates a new policy . Let
The unclipped surrogate is , an importance-sampled policy-gradient objective. PPO replaces it with a pessimistic clipped version [schulman2017proximal]:
The case analysis is the point. If , increasing helps, so PPO follows the gradient only until ; beyond that the clipped constant is smaller and the gradient is zero. If , decreasing helps, so PPO follows the gradient only until ; below that the clipped constant is smaller and the gradient is zero. Away from the kink,
def ppo_clipped_objective_and_grad(
logits,
old_logits,
actions,
advantages,
clip_eps=0.2,
):
"""Mean PPO clipped surrogate and its gradient for one categorical state."""
logits = np.asarray(logits, dtype=np.float64)
old_logits = np.asarray(old_logits, dtype=np.float64)
actions = np.asarray(actions, dtype=np.int64)
advantages = np.asarray(advantages, dtype=np.float64)
probs = softmax(logits)
old_probs = softmax(old_logits)
ratios = probs[actions] / old_probs[actions]
clipped = np.clip(ratios, 1.0 - clip_eps, 1.0 + clip_eps)
objective_terms = np.minimum(ratios * advantages, clipped * advantages)
active = np.where(
advantages >= 0.0,
ratios <= 1.0 + clip_eps,
ratios >= 1.0 - clip_eps,
)
grad = np.zeros_like(logits)
for action, ratio, advantage, is_active in zip(actions, ratios, advantages, active):
if not is_active:
continue
grad_logp = -probs.copy()
grad_logp[action] += 1.0
grad += advantage * ratio * grad_logp
return float(np.mean(objective_terms)), grad / len(actions)
The tests gradient-check the active region and separately assert that the blocked directions have zero gradient.
The min makes the objective pessimistic. It never gives extra credit for moving a sampled action’s probability beyond the trust interval in the helpful direction, but it still penalizes movement in the harmful direction. For , making the action much less likely remains bad and keeps a gradient. For , making the action much more likely remains bad and keeps a gradient. The clip is therefore not symmetric in gradient space; it blocks only the update that already did enough of what the advantage asked.
35.4 Value loss, entropy, and a toy run
PPO implementations usually optimize more than the clipped policy surrogate. A value head learns returns with a squared loss, , so advantages can be estimated instead of using raw returns. An entropy bonus, , discourages premature collapse while the policy is still exploring. The total objective is therefore policy surrogate minus value-loss weight plus entropy weight, with signs chosen for gradient ascent on policy quality.
These extra terms do different jobs. The value loss is supervised regression on returns and is optimized by ordinary backpropagation. The entropy bonus is not a reward-model score; it is an exploration regularizer that keeps the categorical distribution broad enough to keep discovering alternatives. In LLM RLHF, entropy bonuses are often small compared with KL control, but the concept is the same: avoid collapsing the policy before the reward signal has been explored.
The toy run below has one state and three actions with known rewards. It samples actions from the old categorical policy, subtracts the batch mean as a baseline, applies several PPO ascent steps, and repeats. The test checks that the final policy puts more than 90% probability on the known best action.
This toy is deliberately simpler than a text model. There is one state, the value baseline is just the batch mean, and rewards are exact. What remains is the part PPO contributes: the gradient uses likelihood ratios against a frozen old policy, and clipping prevents repeated passes over the same batch from making an unbounded probability jump.
def categorical_entropy(probs):
probs = np.asarray(probs, dtype=np.float64)
return float(-np.sum(probs * np.log(probs)))
def value_loss_and_grad(values, returns):
values = np.asarray(values, dtype=np.float64)
returns = np.asarray(returns, dtype=np.float64)
diff = values - returns
return float(0.5 * np.mean(diff * diff)), diff / diff.size
def toy_ppo_run(action_rewards, steps=80, batch_size=96, lr=0.35, seed=0):
"""PPO on a one-state categorical policy with known action rewards."""
rng = np.random.default_rng(seed)
rewards = np.asarray(action_rewards, dtype=np.float64)
logits = np.zeros_like(rewards)
history = []
for _ in range(steps):
old_logits = logits.copy()
old_probs = softmax(old_logits)
actions = rng.choice(len(rewards), size=batch_size, p=old_probs)
batch_rewards = rewards[actions]
advantages = batch_rewards - np.mean(batch_rewards)
for _ in range(4):
_, grad = ppo_clipped_objective_and_grad(
logits,
old_logits,
actions,
advantages,
)
logits += lr * grad
history.append(softmax(logits))
return logits, np.array(history)
|
In practice
|
The RLHF recipe of preference data, a learned reward model, and policy optimization was used for summarization and instruction-following systems [christiano2017deep] [stiennon2020learning] [ouyang2022training]. Current LLM post-training may replace human comparisons with AI feedback or verifiable rewards, but it still balances reward against drift from a reference policy [lambert2025reinforcement]. PPO remains important because clipping is a simple local trust region: it limits how much one batch can change sampled action probabilities. Reward models are not ground truth; KL penalties, clipping, value losses, and evaluation are all guardrails around that weakness. |
35.5 Teach it
The one-sentence version. RLHF learns a reward from preferences, then PPO nudges the policy toward high-reward samples while clipping steps that move too far.
An analogy. A reward model is a judge trained from pairwise taste tests; PPO is a cautious editor that accepts improvements only in small revisions.
At the board.
-
Write two responses and a sigmoid over .
-
Add the RLHF objective: reward minus times KL to the reference.
-
Draw against the probability ratio for positive and negative .
-
Mark the flat clipped regions and write .
Misconceptions to address.
-
"The reward model is the goal." It is a learned proxy.
-
"PPO clipping clips gradients directly." It clips the objective, which creates zero-gradient regions.
-
"KL and clipping are redundant." KL anchors to a reference; clipping limits one update batch.
Check for understanding. For a positive advantage, why should PPO stop increasing the objective once the new policy makes the sampled action much more likely than the old policy did?
35.6 Exercises
For two responses with rewards and , compute the Bradley-Terry preference probability for the first response and the loss if it is the winner.
Derive the gradient of (35.2) for a linear reward model . Explain why the update raises the winner’s reward and lowers the loser’s reward when the margin is too small.
Analyze (35.5) for and . In which ratio ranges is the gradient active, and where is it zero?
Use the chapter code to run PPO on rewards . Verify that the best action’s probability increases, and check the value-loss gradient on a two-element example.
References
-
[bradley1952] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39, 324–345, 1952.
-
[christiano2017deep] P. Christiano et al. Deep reinforcement learning from human preferences. 2017. arXiv:1706.03741
-
[lambert2025reinforcement] N. Lambert. Reinforcement Learning from Human Feedback. 2025. arXiv:2504.12501
-
[ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155
-
[schulman2017proximal] J. Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347
-
[stiennon2020learning] N. Stiennon et al. Learning to summarize from human feedback. 2020. arXiv:2009.01325