Chapter 36
Direct Preference Optimization
From the KL-regularized optimum to the DPO loss, its gradient, and its variants.
Direct Preference Optimization (DPO) turns pairwise preferences into a supervised loss for a policy, without first fitting a separate reward model and then running PPO. For 2026 LLMs it is the simplest way to say "make the chosen answer more likely than the rejected one, but stay near the reference model." The whole method comes from solving one KL-regularized control problem and substituting that solution into a Bradley—Terry preference model [rafailov2023direct].
The chapter uses one prompt at a time. In a real language model, is a whole answer and is the sum of token log-probabilities under teacher forcing. Treating an answer as one categorical outcome keeps the algebra visible: DPO changes sequence probabilities, not just the final token. The reference policy is frozen, so it acts as a ruler for measuring how far the trainable policy has moved. That ruler matters: a long, generic answer may be likely under both models, while a sharper answer is useful only if the trained policy prefers it more than the reference already did.
36.1 The KL-regularized policy
Fix one prompt, a finite set of answers , a reference policy , a reward , and . The regularized objective is
Let
Then a direct expansion gives
By Gibbs' inequality, reviewed in Section 7.4, the KL term is nonnegative and is zero only at . Thus the optimum tilts the reference toward high-reward answers by an exponential factor. The reference prevents arbitrary drift, and sets how expensive drift is: small sharpens the tilt; large keeps close to .
This derivation is more useful than a Lagrange-multiplier solution because it also tells us the value of being optimal: . The partition function is the reference policy’s moment-generating average of . If every reward is shifted by a constant, is multiplied by the same exponential factor and is unchanged. Only reward differences can affect choices, which is exactly what preference data observes.
def kl_regularized_optimum(reference, reward, beta):
"""The policy pi*(y) proportional to pi_ref(y) exp(r(y) / beta)."""
reference = np.asarray(reference, dtype=np.float64)
reward = np.asarray(reward, dtype=np.float64)
unnormalized = reference * np.exp(reward / beta)
return unnormalized / np.sum(unnormalized)
36.2 The policy is a reward model
Invert (36.2):
This says a policy carries an implicit reward: answers it raises above the reference have positive reward, up to the prompt-dependent constant . In preference data that constant disappears. The Bradley—Terry model says a winner beats a loser with probability [bradley1952]
Substitute the implicit reward of the trainable policy . The two terms cancel, leaving the DPO logit
The loss for one preference pair is just binary cross-entropy with target "winner":
Notice what is and is not fitted. We never learn an absolute reward value, because the unknown would make that impossible from pairwise labels alone. We only learn reward differences that can be represented as policy-reference log-ratios. That is enough for training, since the Bradley—Terry likelihood uses only . It also means the scale of the implicit reward is tied to ; changing changes the logits seen by the preference loss.
36.3 The gradient weight
Differentiate (36.7) with respect to the DPO logit:
Therefore
The weight is large when the current policy still scores the loser near or above the winner, and small when the preference is already satisfied. Thus DPO automatically focuses updates on pairs it is getting wrong. The reference terms shape but do not receive gradients.
For a single softmax over the same answer set, the normalization inside cancels. In a sequence model the two answers visit different token contexts, so the gradient is the usual difference of two teacher-forced log-probability gradients. Either way, the sign is simple: increase the winner and decrease the loser, scaled by . The tests finite-difference the categorical version so the displayed backward pass is not just a symbolic claim.
def dpo_loss_and_grad(logits, reference, pairs, beta):
"""Mean DPO loss and gradient for (winner, loser) categorical pairs."""
logits = np.asarray(logits, dtype=np.float64)
logp = log_softmax(logits)
log_ref = np.log(np.asarray(reference, dtype=np.float64))
grad = np.zeros_like(logits)
losses = []
for winner, loser in pairs:
delta = beta * ((logp[winner] - log_ref[winner])
- (logp[loser] - log_ref[loser]))
weight = sigmoid(-delta)
losses.append(np.logaddexp(0.0, -delta))
grad[winner] -= beta * weight
grad[loser] += beta * weight
return float(np.mean(losses)), grad / len(pairs)
36.4 A tiny DPO run
For a categorical policy over a few answers, the expected DPO objective can be computed over all unordered pairs. If the Bradley—Terry targets come from a reward , the minimum is any logit vector whose log-ratio to the reference differs from by a constant. After softmax, that is exactly (36.2). The test for this chapter gradient-checks the loss and verifies that the optimized categorical policy matches the closed-form optimum.
This toy run is deliberately small, but it captures the consistency story. The synthetic preference probabilities are generated from a known reward, the model sees every pair, and the optimizer is allowed to represent any categorical distribution. Under those conditions the DPO minimum recovers the same policy as KL-regularized reward maximization. Real data violates all three assumptions: preferences are sampled sparsely, labels can be inconsistent, and the model shares parameters across many prompts. The derivation still explains what the loss is trying to estimate.
def fit_toy_dpo(reference, rewards, beta=0.7, steps=800, lr=0.8):
"""Optimize a tiny categorical policy and return (policy, optimum)."""
logits = np.log(np.asarray(reference, dtype=np.float64))
for _ in range(steps):
_, grad = soft_dpo_loss_and_grad(logits, reference, rewards, beta)
logits -= lr * grad
policy = np.exp(log_softmax(logits))
optimum = kl_regularized_optimum(reference, rewards, beta)
return policy, optimum
|
In practice
|
DPO is popular because it reuses the supervised fine-tuning pipeline: batches contain a prompt, a chosen answer, and a rejected answer, and the loss uses only log-probabilities from the policy and frozen reference [rafailov2023direct]. The reference is usually the SFT model or a copy of the policy before preference training. The method inherits the Bradley—Terry assumption from classical paired-comparison models [bradley1952], so noisy or inconsistent preferences become noisy labels rather than a separately inspected reward model. Practitioners tune as a behavior knob: too small can overfit preference artifacts, while too large leaves the policy close to the reference. This chapter omits IPO, KTO, and SimPO because the shared bibliography for this book does not yet contain their sources. |
36.5 Teach it
DPO is preference learning after eliminating the reward model algebraically. Analogy: instead of building a thermometer for reward and then steering by it, compare the policy to the reference and use that log-ratio as the thermometer. Board steps: write ; complete the square into ; invert to get ; subtract rewards in Bradley—Terry so cancels. Misconceptions: DPO is not unregularized supervised learning, because the reference ratio is the reward scale; is not a learning rate, because it changes the preference logit; a saturated preference pair gives little gradient. Check: when the model already makes the winner much more likely than the loser relative to the reference, should the DPO update on that pair be large or small?
36.6 Exercises
Use (36.4) inside the Bradley—Terry model and show why cannot affect a preference between two answers to the same prompt.
Differentiate the DPO loss and explain why emphasizes hard or wrong preference pairs.
Implement expected DPO for all unordered pairs of a categorical policy and verify that gradient descent reaches .
References
-
[kullback1951] S. Kullback and R. A. Leibler. On information and sufficiency. Annals of Mathematical Statistics 22(1), 79–86, 1951.
-
[bradley1952] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39, 324–345, 1952.
-
[rafailov2023direct] R. Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. 2023. arXiv:2305.18290