Chapter 35

Reward Models, PPO & RLHF

Bradley-Terry reward models, the clipped PPO objective, and KL-regularized RLHF.

RLHF turns human or AI preferences into a reward, then uses reinforcement learning to move a language-model policy toward high-reward responses without drifting too far from a reference model. The pipeline is three pieces: learn a reward model from comparisons, define a KL-regularized objective, and optimize the policy with a stable policy-gradient method. PPO is the standard small step in that last piece.

35.1 Reward models from pairwise preferences

A preference dataset contains pairs: for the same prompt, response ywy_w was preferred to response yly_l. A Bradley-Terry model says the probability of that preference is a sigmoid of reward difference [bradley1952]:

P(yw≻yl)=σ(rϕ(yw)−rϕ(yl)).(35.1)P(y_w \succ y_l) = \sigma(r_\vphi(y_w)-r_\vphi(y_l)).\tag{35.1}

For one pair with d=rw−rld=r_w-r_l, the negative log-likelihood is

ℓ(d)=−log⁡σ(d)=log⁡(1+e−d).(35.2)\ell(d) = -\log\sigma(d) = \log(1+e^{-d}).\tag{35.2}

Differentiating gives ∂ℓ/∂d=σ(d)−1\partial\ell/\partial d=\sigma(d)-1. For a linear reward rϕ(y)=xy⊤ϕr_\vphi(y)=\vx_y^\T\vphi, the gradient is (σ(d)−1)(xw−xl)(\sigma(d)-1)(\vx_w-\vx_l). The code builds synthetic pairs from a hidden linear reward, gradient-checks this loss, and trains a small reward model to score winners above losers.

Only differences matter. Adding the same constant to both rewards leaves dd unchanged, so a reward model’s absolute zero is arbitrary. That is fine for policy optimization, which needs to rank sampled responses, but it means reward values should not be read like calibrated human scores. The synthetic data keeps one prompt implicit and uses feature vectors for responses; real systems condition the reward model on both prompt and response.

Listing 35.1 Bradley-Terry reward model
def bradley_terry_loss_and_grad(weights, winners, losers):
    """Loss -log sigmoid(r_w - r_l) for a linear reward model."""
    weights = np.asarray(weights, dtype=np.float64)
    features = np.asarray(winners) - np.asarray(losers)
    margins = features @ weights
    loss = np.mean(np.logaddexp(0.0, -margins))
    sigmoid = 1.0 / (1.0 + np.exp(-margins))
    grad = ((sigmoid - 1.0)[:, None] * features).mean(axis=0)
    return float(loss), grad


def train_reward_model(winners, losers, steps=300, lr=0.5):
    weights = np.zeros(winners.shape[1], dtype=np.float64)
    for _ in range(steps):
        _, grad = bradley_terry_loss_and_grad(weights, winners, losers)
        weights -= lr * grad
    return weights

35.2 KL-regularized RLHF

Once a reward model exists, the policy objective for a prompt xx is

J(θ)=Ey∼πθ(⋅∣x)[rϕ(x,y)]−βDKL(πθ(⋅∣x)∥πref(⋅∣x)).(35.3)J(\vtheta) = \E_{y\sim\pi_\vtheta(\cdot\mid x)}[r_\vphi(x,y)] - \beta\KL(\pi_\vtheta(\cdot\mid x)\Vert\pi_{ref}(\cdot\mid x)).\tag{35.3}

The reward term pulls the policy toward preferred responses. The KL term keeps it close to a reference policy, usually the SFT model from Chapter 33, so reward-model mistakes do not dominate. The coefficient β\beta is the exchange rate between reward and drift.

The KL direction matters. Because responses are sampled from the current policy, the penalty is naturally estimated under πθ\pi_\vtheta, giving DKL(πθ∥πref)\KL(\pi_\vtheta\Vert\pi_{ref}). This is the same reverse direction discussed in Section 7.4.1: it punishes probability mass the new policy puts where the reference was unlikely. A large β\beta keeps style and coverage close to the reference; a small β\beta lets the reward model steer more aggressively.

The full response-level KL sums over impossible-to-enumerate outputs. In practice, algorithms estimate a per-token KL on sampled responses; Section 7.8 derives common estimators and their variance behavior. The tiny categorical code computes the exact KL so the objective can be tested without sampling noise.

This objective is not yet an algorithm. It says which expectation to improve, but a language model cannot enumerate every response and take an exact gradient. PPO supplies the sampled, local update rule: collect completions from the current policy, freeze that policy as πold\pi_{old}, estimate advantages, then take a few cautious gradient steps on the same batch.

Listing 35.2 KL-regularized categorical objective
def categorical_kl(policy_probs, reference_probs):
    """KL(pi || pi_ref) for categorical policies."""
    policy_probs = np.asarray(policy_probs, dtype=np.float64)
    reference_probs = np.asarray(reference_probs, dtype=np.float64)
    log_ratio = np.log(policy_probs) - np.log(reference_probs)
    return float(np.sum(policy_probs * log_ratio))


def rlhf_objective(policy_probs, rewards, reference_probs, beta):
    """E_pi[r] - beta KL(pi || pi_ref)."""
    reward_term = float(np.asarray(policy_probs) @ np.asarray(rewards))
    return reward_term - beta * categorical_kl(policy_probs, reference_probs)

35.3 PPO’s clipped surrogate

A policy-gradient update uses samples from an old policy πold\pi_{old} but evaluates a new policy πθ\pi_\vtheta. Let

rt(θ)=πθ(at∣st)πold(at∣st).(35.4)r_t(\vtheta)=\frac{\pi_\vtheta(a_t\mid s_t)}{\pi_{old}(a_t\mid s_t)}.\tag{35.4}

The unclipped surrogate is rtAtr_t A_t, an importance-sampled policy-gradient objective. PPO replaces it with a pessimistic clipped version [schulman2017proximal]:

LtCLIP(θ)=min⁡(rtAt,clip⁡(rt,1−ϵ,1+ϵ)At).(35.5)L^{CLIP}_t(\vtheta)=\min\big(r_t A_t, \operatorname{clip}(r_t,1-\epsilon,1+\epsilon)A_t\big).\tag{35.5}

The case analysis is the point. If At>0A_t>0, increasing rtr_t helps, so PPO follows the gradient only until rt>1+ϵr_t>1+\epsilon; beyond that the clipped constant is smaller and the gradient is zero. If At<0A_t<0, decreasing rtr_t helps, so PPO follows the gradient only until rt<1−ϵr_t<1-\epsilon; below that the clipped constant is smaller and the gradient is zero. Away from the kink,

∇θ(rtAt)=Atrt∇θlog⁡πθ(at∣st).(35.6)\nabla_\vtheta(r_t A_t)=A_t r_t\nabla_\vtheta\log\pi_\vtheta(a_t\mid s_t).\tag{35.6}
Listing 35.3 PPO clipped surrogate and gradient
def ppo_clipped_objective_and_grad(
    logits,
    old_logits,
    actions,
    advantages,
    clip_eps=0.2,
):
    """Mean PPO clipped surrogate and its gradient for one categorical state."""
    logits = np.asarray(logits, dtype=np.float64)
    old_logits = np.asarray(old_logits, dtype=np.float64)
    actions = np.asarray(actions, dtype=np.int64)
    advantages = np.asarray(advantages, dtype=np.float64)
    probs = softmax(logits)
    old_probs = softmax(old_logits)
    ratios = probs[actions] / old_probs[actions]
    clipped = np.clip(ratios, 1.0 - clip_eps, 1.0 + clip_eps)
    objective_terms = np.minimum(ratios * advantages, clipped * advantages)

    active = np.where(
        advantages >= 0.0,
        ratios <= 1.0 + clip_eps,
        ratios >= 1.0 - clip_eps,
    )
    grad = np.zeros_like(logits)
    for action, ratio, advantage, is_active in zip(actions, ratios, advantages, active):
        if not is_active:
            continue
        grad_logp = -probs.copy()
        grad_logp[action] += 1.0
        grad += advantage * ratio * grad_logp
    return float(np.mean(objective_terms)), grad / len(actions)

The tests gradient-check the active region and separately assert that the blocked directions have zero gradient.

The min makes the objective pessimistic. It never gives extra credit for moving a sampled action’s probability beyond the trust interval in the helpful direction, but it still penalizes movement in the harmful direction. For At>0A_t>0, making the action much less likely remains bad and keeps a gradient. For At<0A_t<0, making the action much more likely remains bad and keeps a gradient. The clip is therefore not symmetric in gradient space; it blocks only the update that already did enough of what the advantage asked.

35.4 Value loss, entropy, and a toy run

PPO implementations usually optimize more than the clipped policy surrogate. A value head learns returns with a squared loss, 12(V(st)−Gt)2\tfrac12(V(s_t)-G_t)^2, so advantages can be estimated instead of using raw returns. An entropy bonus, cHH(πθ(⋅∣st))c_H H(\pi_\vtheta(\cdot\mid s_t)), discourages premature collapse while the policy is still exploring. The total objective is therefore policy surrogate minus value-loss weight plus entropy weight, with signs chosen for gradient ascent on policy quality.

These extra terms do different jobs. The value loss is supervised regression on returns and is optimized by ordinary backpropagation. The entropy bonus is not a reward-model score; it is an exploration regularizer that keeps the categorical distribution broad enough to keep discovering alternatives. In LLM RLHF, entropy bonuses are often small compared with KL control, but the concept is the same: avoid collapsing the policy before the reward signal has been explored.

The toy run below has one state and three actions with known rewards. It samples actions from the old categorical policy, subtracts the batch mean as a baseline, applies several PPO ascent steps, and repeats. The test checks that the final policy puts more than 90% probability on the known best action.

This toy is deliberately simpler than a text model. There is one state, the value baseline is just the batch mean, and rewards are exact. What remains is the part PPO contributes: the gradient uses likelihood ratios against a frozen old policy, and clipping prevents repeated passes over the same batch from making an unbounded probability jump.

Listing 35.4 Value loss, entropy, and toy PPO
def categorical_entropy(probs):
    probs = np.asarray(probs, dtype=np.float64)
    return float(-np.sum(probs * np.log(probs)))


def value_loss_and_grad(values, returns):
    values = np.asarray(values, dtype=np.float64)
    returns = np.asarray(returns, dtype=np.float64)
    diff = values - returns
    return float(0.5 * np.mean(diff * diff)), diff / diff.size


def toy_ppo_run(action_rewards, steps=80, batch_size=96, lr=0.35, seed=0):
    """PPO on a one-state categorical policy with known action rewards."""
    rng = np.random.default_rng(seed)
    rewards = np.asarray(action_rewards, dtype=np.float64)
    logits = np.zeros_like(rewards)
    history = []
    for _ in range(steps):
        old_logits = logits.copy()
        old_probs = softmax(old_logits)
        actions = rng.choice(len(rewards), size=batch_size, p=old_probs)
        batch_rewards = rewards[actions]
        advantages = batch_rewards - np.mean(batch_rewards)
        for _ in range(4):
            _, grad = ppo_clipped_objective_and_grad(
                logits,
                old_logits,
                actions,
                advantages,
            )
            logits += lr * grad
        history.append(softmax(logits))
    return logits, np.array(history)
In practice

The RLHF recipe of preference data, a learned reward model, and policy optimization was used for summarization and instruction-following systems [christiano2017deep] [stiennon2020learning] [ouyang2022training]. Current LLM post-training may replace human comparisons with AI feedback or verifiable rewards, but it still balances reward against drift from a reference policy [lambert2025reinforcement]. PPO remains important because clipping is a simple local trust region: it limits how much one batch can change sampled action probabilities. Reward models are not ground truth; KL penalties, clipping, value losses, and evaluation are all guardrails around that weakness.

Key equations
P(yw≻yl)=σ(rϕ(yw)−rϕ(yl)),ℓ(d)=−log⁡σ(d)P(y_w \succ y_l)=\sigma(r_\vphi(y_w)-r_\vphi(y_l)), \qquad \ell(d)=-\log\sigma(d)
J(θ)=Ey∼πθ[rϕ(x,y)]−βDKL(πθ(⋅∣x)∥πref(⋅∣x))J(\vtheta)=\E_{y\sim\pi_\vtheta}[r_\vphi(x,y)] -\beta\KL(\pi_\vtheta(\cdot\mid x)\Vert\pi_{ref}(\cdot\mid x))
rt(θ)=πθ(at∣st)πold(at∣st)r_t(\vtheta)=\frac{\pi_\vtheta(a_t\mid s_t)}{\pi_{old}(a_t\mid s_t)}
LtCLIP=min⁡(rtAt,clip⁡(rt,1−ϵ,1+ϵ)At)L^{CLIP}_t=\min\big(r_tA_t, \operatorname{clip}(r_t,1-\epsilon,1+\epsilon)A_t\big)
∇(rtAt)=Atrt∇log⁡πθ(at∣st)\nabla(r_tA_t)=A_t r_t\nabla\log\pi_\vtheta(a_t\mid s_t)

35.5 Teach it

The one-sentence version. RLHF learns a reward from preferences, then PPO nudges the policy toward high-reward samples while clipping steps that move too far.

An analogy. A reward model is a judge trained from pairwise taste tests; PPO is a cautious editor that accepts improvements only in small revisions.

At the board.

  1. Write two responses and a sigmoid over rw−rlr_w-r_l.

  2. Add the RLHF objective: reward minus β\beta times KL to the reference.

  3. Draw rArA against the probability ratio for positive and negative AA.

  4. Mark the flat clipped regions and write Ar∇log⁡πA r\nabla\log\pi.

Misconceptions to address.

  • "The reward model is the goal." It is a learned proxy.

  • "PPO clipping clips gradients directly." It clips the objective, which creates zero-gradient regions.

  • "KL and clipping are redundant." KL anchors to a reference; clipping limits one update batch.

Check for understanding. For a positive advantage, why should PPO stop increasing the objective once the new policy makes the sampled action much more likely than the old policy did?

35.6 Exercises

Exercise 35.1 ★ Preference likelihood

For two responses with rewards 22 and 0.50.5, compute the Bradley-Terry preference probability for the first response and the loss if it is the winner.

Exercise 35.2 ★★ Reward-model gradient

Derive the gradient of (35.2) for a linear reward model rϕ(y)=xy⊤ϕr_\vphi(y)=\vx_y^\T\vphi. Explain why the update raises the winner’s reward and lowers the loser’s reward when the margin is too small.

Exercise 35.3 ★★ PPO clipping cases

Analyze (35.5) for A>0A>0 and A<0A<0. In which ratio ranges is the gradient Ar∇log⁡πA r\nabla\log\pi active, and where is it zero?

Exercise 35.4 ★★★ Toy PPO implementation

Use the chapter code to run PPO on rewards [0,1,−0.2][0,1,-0.2]. Verify that the best action’s probability increases, and check the value-loss gradient on a two-element example.

References

  • [bradley1952] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39, 324–345, 1952.

  • [christiano2017deep] P. Christiano et al. Deep reinforcement learning from human preferences. 2017. arXiv:1706.03741

  • [lambert2025reinforcement] N. Lambert. Reinforcement Learning from Human Feedback. 2025. arXiv:2504.12501

  • [ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155

  • [schulman2017proximal] J. Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347

  • [stiennon2020learning] N. Stiennon et al. Learning to summarize from human feedback. 2020. arXiv:2009.01325