Chapter 36

Direct Preference Optimization

From the KL-regularized optimum to the DPO loss, its gradient, and its variants.

Direct Preference Optimization (DPO) turns pairwise preferences into a supervised loss for a policy, without first fitting a separate reward model and then running PPO. For 2026 LLMs it is the simplest way to say "make the chosen answer more likely than the rejected one, but stay near the reference model." The whole method comes from solving one KL-regularized control problem and substituting that solution into a Bradley—​Terry preference model [rafailov2023direct].

The chapter uses one prompt at a time. In a real language model, yy is a whole answer and log⁡πθ(y)\log \pi_\theta(y) is the sum of token log-probabilities under teacher forcing. Treating an answer as one categorical outcome keeps the algebra visible: DPO changes sequence probabilities, not just the final token. The reference policy is frozen, so it acts as a ruler for measuring how far the trainable policy has moved. That ruler matters: a long, generic answer may be likely under both models, while a sharper answer is useful only if the trained policy prefers it more than the reference already did.

36.1 The KL-regularized policy

Fix one prompt, a finite set of answers yy, a reference policy π0\pi_0, a reward r(y)r(y), and β>0\beta > 0. The regularized objective is

J(π)=Ey∼π[r(y)]−β DKL(π ∥ π0).(36.1)J(\pi) = \E_{y \sim \pi}[r(y)] - \beta\,\KL(\pi\,\Vert\,\pi_0) .\tag{36.1}

Let

π∗(y)=π0(y)exp⁡(r(y)/β)Z,Z=∑yπ0(y)exp⁡(r(y)/β).(36.2)\pi^*(y) = \frac{\pi_0(y)\exp(r(y)/\beta)}{Z}, \quad Z = \sum_y \pi_0(y)\exp(r(y)/\beta).\tag{36.2}

Then a direct expansion gives

J(π)=β∑yπ(y)log⁡π0(y)er(y)/βπ(y)=βlog⁡Z−βDKL(π ∥ π∗).(36.3)\begin{aligned} J(\pi) &= \beta \sum_y \pi(y)\log\frac{\pi_0(y)e^{r(y)/\beta}}{\pi(y)} \\ &= \beta\log Z - \beta\KL(\pi\,\Vert\,\pi^*). \end{aligned}\tag{36.3}

By Gibbs' inequality, reviewed in Section 7.4, the KL term is nonnegative and is zero only at π=π∗\pi = \pi^*. Thus the optimum tilts the reference toward high-reward answers by an exponential factor. The reference prevents arbitrary drift, and β\beta sets how expensive drift is: small β\beta sharpens the tilt; large β\beta keeps π∗\pi^* close to π0\pi_0.

This derivation is more useful than a Lagrange-multiplier solution because it also tells us the value of being optimal: βlog⁡Z\beta\log Z. The partition function ZZ is the reference policy’s moment-generating average of exp⁡(r/β)\exp(r/\beta). If every reward is shifted by a constant, ZZ is multiplied by the same exponential factor and π∗\pi^* is unchanged. Only reward differences can affect choices, which is exactly what preference data observes.

Listing 36.1 The closed-form optimum of the KL-regularized objective
def kl_regularized_optimum(reference, reward, beta):
    """The policy pi*(y) proportional to pi_ref(y) exp(r(y) / beta)."""
    reference = np.asarray(reference, dtype=np.float64)
    reward = np.asarray(reward, dtype=np.float64)
    unnormalized = reference * np.exp(reward / beta)
    return unnormalized / np.sum(unnormalized)

36.2 The policy is a reward model

Invert (36.2):

r(y)=βlog⁡π∗(y)π0(y)+βlog⁡Z.(36.4)r(y) = \beta\log\frac{\pi^*(y)}{\pi_0(y)} + \beta\log Z .\tag{36.4}

This says a policy carries an implicit reward: answers it raises above the reference have positive reward, up to the prompt-dependent constant βlog⁡Z\beta\log Z. In preference data that constant disappears. The Bradley—​Terry model says a winner ywy_w beats a loser yly_l with probability [bradley1952]

P(yw≻yl)=σ(r(yw)−r(yl)).(36.5)P(y_w \succ y_l) = \sigma\big(r(y_w) - r(y_l)\big).\tag{36.5}

Substitute the implicit reward of the trainable policy πθ\pi_\theta. The two βlog⁡Z\beta\log Z terms cancel, leaving the DPO logit

δ^θ=β(log⁡πθ(yw)π0(yw)−log⁡πθ(yl)π0(yl)).(36.6)\hat{\delta}_\theta = \beta\Big( \log\frac{\pi_\theta(y_w)}{\pi_0(y_w)} - \log\frac{\pi_\theta(y_l)}{\pi_0(y_l)}\Big).\tag{36.6}

The loss for one preference pair is just binary cross-entropy with target "winner":

LDPO=−log⁡σ(δ^θ).(36.7)\mathcal{L}_{\mathrm{DPO}} = -\log \sigma(\hat{\delta}_\theta).\tag{36.7}

Notice what is and is not fitted. We never learn an absolute reward value, because the unknown ZZ would make that impossible from pairwise labels alone. We only learn reward differences that can be represented as policy-reference log-ratios. That is enough for training, since the Bradley—​Terry likelihood uses only rw−rlr_w-r_l. It also means the scale of the implicit reward is tied to β\beta; changing β\beta changes the logits seen by the preference loss.

36.3 The gradient weight

Differentiate (36.7) with respect to the DPO logit:

∂L∂δ^=σ(δ^)−1=−σ(−δ^).(36.8)\frac{\partial \mathcal{L}}{\partial \hat{\delta}} = \sigma(\hat{\delta}) - 1 = -\sigma(-\hat{\delta}).\tag{36.8}

Therefore

∇θL=−βσ(−δ^θ)∇θ(log⁡πθ(yw)−log⁡πθ(yl)).(36.9)\nabla_\theta \mathcal{L} = -\beta\sigma(-\hat{\delta}_\theta) \nabla_\theta\big(\log\pi_\theta(y_w)-\log\pi_\theta(y_l)\big).\tag{36.9}

The weight σ(−δ^)=σ(r^l−r^w)\sigma(-\hat{\delta}) = \sigma(\hat{r}_l - \hat{r}_w) is large when the current policy still scores the loser near or above the winner, and small when the preference is already satisfied. Thus DPO automatically focuses updates on pairs it is getting wrong. The reference terms shape δ^\hat{\delta} but do not receive gradients.

For a single softmax over the same answer set, the normalization inside log⁡πθ(yw)−log⁡πθ(yl)\log\pi_\theta(y_w)-\log\pi_\theta(y_l) cancels. In a sequence model the two answers visit different token contexts, so the gradient is the usual difference of two teacher-forced log-probability gradients. Either way, the sign is simple: increase the winner and decrease the loser, scaled by βσ(r^l−r^w)\beta\sigma(\hat{r}_l-\hat{r}_w). The tests finite-difference the categorical version so the displayed backward pass is not just a symbolic claim.

Listing 36.2 DPO loss and analytic gradient for a categorical policy
def dpo_loss_and_grad(logits, reference, pairs, beta):
    """Mean DPO loss and gradient for (winner, loser) categorical pairs."""
    logits = np.asarray(logits, dtype=np.float64)
    logp = log_softmax(logits)
    log_ref = np.log(np.asarray(reference, dtype=np.float64))
    grad = np.zeros_like(logits)
    losses = []
    for winner, loser in pairs:
        delta = beta * ((logp[winner] - log_ref[winner])
                        - (logp[loser] - log_ref[loser]))
        weight = sigmoid(-delta)
        losses.append(np.logaddexp(0.0, -delta))
        grad[winner] -= beta * weight
        grad[loser] += beta * weight
    return float(np.mean(losses)), grad / len(pairs)

36.4 A tiny DPO run

For a categorical policy over a few answers, the expected DPO objective can be computed over all unordered pairs. If the Bradley—​Terry targets come from a reward rr, the minimum is any logit vector whose log-ratio to the reference differs from r/βr/\beta by a constant. After softmax, that is exactly (36.2). The test for this chapter gradient-checks the loss and verifies that the optimized categorical policy matches the closed-form optimum.

This toy run is deliberately small, but it captures the consistency story. The synthetic preference probabilities are generated from a known reward, the model sees every pair, and the optimizer is allowed to represent any categorical distribution. Under those conditions the DPO minimum recovers the same policy as KL-regularized reward maximization. Real data violates all three assumptions: preferences are sampled sparsely, labels can be inconsistent, and the model shares parameters across many prompts. The derivation still explains what the loss is trying to estimate.

Listing 36.3 Optimizing DPO on all preference pairs
def fit_toy_dpo(reference, rewards, beta=0.7, steps=800, lr=0.8):
    """Optimize a tiny categorical policy and return (policy, optimum)."""
    logits = np.log(np.asarray(reference, dtype=np.float64))
    for _ in range(steps):
        _, grad = soft_dpo_loss_and_grad(logits, reference, rewards, beta)
        logits -= lr * grad
    policy = np.exp(log_softmax(logits))
    optimum = kl_regularized_optimum(reference, rewards, beta)
    return policy, optimum
In practice

DPO is popular because it reuses the supervised fine-tuning pipeline: batches contain a prompt, a chosen answer, and a rejected answer, and the loss uses only log-probabilities from the policy and frozen reference [rafailov2023direct]. The reference is usually the SFT model or a copy of the policy before preference training. The method inherits the Bradley—​Terry assumption from classical paired-comparison models [bradley1952], so noisy or inconsistent preferences become noisy labels rather than a separately inspected reward model. Practitioners tune β\beta as a behavior knob: too small can overfit preference artifacts, while too large leaves the policy close to the reference. This chapter omits IPO, KTO, and SimPO because the shared bibliography for this book does not yet contain their sources.

Key equations
J(π)=Eπ[r]−βDKL(π ∥ π0)J(\pi) = \E_\pi[r] - \beta\KL(\pi\,\Vert\,\pi_0)
π∗(y)=π0(y)er(y)/β/Z\pi^*(y) = \pi_0(y)e^{r(y)/\beta}/Z
r^θ(y)=βlog⁡πθ(y)π0(y)+βlog⁡Z\hat{r}_\theta(y) = \beta\log\frac{\pi_\theta(y)}{\pi_0(y)} + \beta\log Z
LDPO=−log⁡σ(r^w−r^l)\mathcal{L}_{\mathrm{DPO}} = -\log\sigma(\hat{r}_w - \hat{r}_l)
∇L=−βσ(r^l−r^w)∇(log⁡πw−log⁡πl)\nabla\mathcal{L} = -\beta\sigma(\hat{r}_l-\hat{r}_w) \nabla(\log\pi_w-\log\pi_l)

36.5 Teach it

DPO is preference learning after eliminating the reward model algebraically. Analogy: instead of building a thermometer for reward and then steering by it, compare the policy to the reference and use that log-ratio as the thermometer. Board steps: write E[r]−βDKL(π∥π0)\E[r]-\beta\KL(\pi\Vert\pi_0); complete the square into βlog⁡Z−βDKL(π∥π∗)\beta\log Z-\beta\KL(\pi\Vert\pi^*); invert to get r=βlog⁡(π/π0)+βlog⁡Zr=\beta\log(\pi/\pi_0)+\beta\log Z; subtract rewards in Bradley—​Terry so ZZ cancels. Misconceptions: DPO is not unregularized supervised learning, because the reference ratio is the reward scale; β\beta is not a learning rate, because it changes the preference logit; a saturated preference pair gives little gradient. Check: when the model already makes the winner much more likely than the loser relative to the reference, should the DPO update on that pair be large or small?

36.6 Exercises

Exercise 36.1 ★ Derive the optimum

Starting from (36.1), derive (36.3) and explain exactly where Gibbs' inequality proves optimality.

Exercise 36.2 ★★ Cancel the partition function

Use (36.4) inside the Bradley—​Terry model and show why ZZ cannot affect a preference between two answers to the same prompt.

Exercise 36.3 ★★ Interpret the gradient

Differentiate the DPO loss and explain why σ(r^l−r^w)\sigma(\hat{r}_l-\hat{r}_w) emphasizes hard or wrong preference pairs.

Exercise 36.4 ★★★ Implement the categorical toy

Implement expected DPO for all unordered pairs of a categorical policy and verify that gradient descent reaches π∗(y)∝π0(y)exp⁡(r(y)/β)\pi^*(y) \propto \pi_0(y)\exp(r(y)/\beta).

References

  • [kullback1951] S. Kullback and R. A. Leibler. On information and sufficiency. Annals of Mathematical Statistics 22(1), 79–86, 1951.

  • [bradley1952] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39, 324–345, 1952.

  • [rafailov2023direct] R. Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. 2023. arXiv:2305.18290