Chapter 38
Distillation & Reasoning Models
Forward and on-policy distillation, test-time compute, and process rewards.
Distillation trains a smaller or cheaper student model to match a stronger teacher. For 2026 LLMs it is how expensive reasoning, search, and reward-guided behavior are turned back into a model that can run in one forward pass. The same equations also explain why test-time compute helps: extra samples can improve answers, and distillation tries to amortize that improvement into the weights. The goal is not to copy the teacher’s parameters. It is to copy a distribution over useful outputs: class probabilities, next-token probabilities, full reasoning traces, or selected solutions from a search procedure.
38.1 Soft targets and temperature
Let teacher logits be , student logits , and temperature . Define softened distributions
Knowledge distillation minimizes the teacher-to-student KL, usually implemented as cross-entropy because the teacher entropy is constant [hinton2015distilling]:
The temperature reveals dark knowledge: a class that is not the teacher’s top choice can still be more plausible than other wrong classes. Different wrong-token probabilities tell the student about similarity structure that a one-hot label hides. For language models, the classes are tokens at each position. A high temperature makes the target less certain, so rare but plausible continuations receive visible probability. During training the student still sees the same context as the teacher; only the target distribution changes.
For the gradient, first ignore the constant teacher entropy. The derivative of with respect to a student logit is , because the softmax saw . Multiplying the loss by gives
At high temperature, itself shrinks like , so the unscaled gradient would shrink like . The multiplier keeps the gradient scale comparable while using soft targets. This scaling is why a distillation loss can be mixed with an ordinary next-token loss without the distillation term disappearing merely because a larger was chosen. It does not make all temperatures equivalent: the target distribution is still softer at larger .
def distillation_loss_and_grad(student_logits, teacher_logits, temperature=1.0,
scale_t2=True):
"""KL teacher_T || student_T, with optional T^2 multiplier."""
student_logits = np.asarray(student_logits, dtype=np.float64)
teacher_logits = np.asarray(teacher_logits, dtype=np.float64)
teacher = softmax(teacher_logits / temperature)
log_student = log_softmax(student_logits / temperature)
loss = -np.sum(teacher * log_student)
grad = (softmax(student_logits / temperature) - teacher) / temperature
if scale_t2:
loss *= temperature ** 2
grad *= temperature ** 2
return float(loss), grad
38.2 Forward and on-policy distillation
Classical distillation is a forward KL: sample or enumerate from the teacher and make the student cover the teacher’s distribution, . It is mass-covering in the same sense as Section 7.4.1: missing a teacher-supported answer is expensive.
On-policy distillation instead samples from the student and compares those samples with the teacher, minimizing a reverse direction such as [agarwal2023onpolicy]. That focuses training on the student’s own mistakes and can be mode-seeking: probability mass the student never samples contributes little. It is useful when the student is already deployed for sampling, or when teacher calls are expensive and should be spent where the student actually goes. Forward KL asks the student to cover what the teacher might say; reverse KL asks whether the student’s own samples are defensible under the teacher. In practice, the two directions answer different data-collection questions. Teacher-sampled data is easy to cache and train like supervised examples. Student-sampled data must be regenerated as the student changes, but it targets errors that the current student actually makes.
def forward_kl_teacher_to_student(teacher_probs, student_probs):
"""KL(teacher || student): mass-covering and teacher-sampled."""
teacher_probs = np.asarray(teacher_probs, dtype=np.float64)
student_probs = np.asarray(student_probs, dtype=np.float64)
return float(np.sum(teacher_probs * (np.log(teacher_probs) -
np.log(student_probs))))
def reverse_kl_student_to_teacher(student_probs, teacher_probs):
"""KL(student || teacher): on-policy and mode-seeking."""
student_probs = np.asarray(student_probs, dtype=np.float64)
teacher_probs = np.asarray(teacher_probs, dtype=np.float64)
return float(np.sum(student_probs * (np.log(student_probs) -
np.log(teacher_probs))))
38.3 Distilling reasoning traces
Reasoning models often spend extra tokens exploring a solution before giving the final answer. DeepSeek-R1 reports using reasoning data from a stronger model to distill reasoning behavior into smaller open models [deepseekai2025deepseekr1]. The student is not only matching a final label; it is trained on traces that show intermediate decomposition, checking, and correction. This is amortization: do expensive search, RL, or sampling once, then train a cheaper model to imitate the resulting behavior.
There is a boundary. Distilling a flawed trace can teach the flaw, and a small student may imitate surface form without preserving the computation that made the trace useful. That is why trace quality, filtering, and final-answer verification remain part of the pipeline. Trace distillation also separates two products of a reasoning run. The final answer can be checked or compared, while the path can teach a format for decomposition. A good student should learn both when to write useful intermediate state and when to stop.
38.4 Test-time compute
If one sample is correct with independent probability , best-of- with a perfect verifier succeeds when at least one sample is correct:
Majority voting, or self-consistency, succeeds when more than half the samples are correct [wang2022selfconsistency]:
Best-of- needs a verifier or reward model to choose among answers; majority voting needs answers that can be canonicalized. Scaling test-time compute studies how to spend such samples rather than only scaling parameters [snell2024scaling]. The binomial formulas are optimistic because samples from one model are correlated. If every sample makes the same mistake, voting cannot help. They are still the right first calculation: they show what independent diversity would buy before accounting for selector quality, correlation, and cost.
def best_of_n_accuracy(p, n):
"""Probability that at least one of n independent samples is correct."""
return 1.0 - (1.0 - p) ** n
def majority_vote_accuracy(p, n):
"""Strict-majority accuracy for n independent samples, each correct with prob p."""
threshold = n // 2 + 1
total = 0.0
for k in range(threshold, n + 1):
total += math.comb(n, k) * p ** k * (1.0 - p) ** (n - k)
return total
38.5 Process and outcome rewards
An outcome reward model scores the final answer; it is easy to pair with a verifier but gives sparse credit. A process reward model scores intermediate reasoning steps, giving denser guidance for search or training; step-level supervision was studied in Let’s Verify Step by Step [lightman2023let]. DeepSeek-R1 emphasizes rule-based outcome rewards for reasoning RL where final answers can be checked [deepseekai2025deepseekr1]. In practice, process rewards can guide how to reason, while outcome rewards judge whether the reasoning got somewhere true. When traces are distilled, these rewards decide which traces become examples. Outcome filtering may keep only successful solutions; process filtering can prefer traces whose intermediate steps are locally valid even before the final answer is known.
|
In practice
|
Distillation shows up after expensive teachers, RLVR runs, rejection sampling, and search. A common pattern is to generate many candidate traces, filter or rank them with outcome or process rewards, and train the next model on the selected traces. Temperature distillation is still used for logits when teacher probabilities are available, but many reasoning pipelines distill text traces instead. The validation question is always the same: did the student learn the capability, or only mimic the teacher’s style? |
38.6 Teach it
Distillation turns expensive behavior into a cheaper student’s probabilities or traces. Analogy: a master solves problems with scratch work; the apprentice studies both the answers and the hints about which wrong answers were close. Board steps: soften teacher and student with temperature; minimize ; choose forward KL for teacher-sampled coverage or reverse KL for student-sampled correction; use binomial sums to price extra samples at test time. Misconceptions: temperature is not randomness during training, it changes target probabilities; best-of- assumes a selector; majority voting helps only if samples are more likely right than wrong and errors are not perfectly correlated. Check: why does appear in the loss?
38.7 Exercises
Given teacher logits , describe how increasing changes the target probabilities and why the smaller nonzero probabilities can help a student.
Starting from cross-entropy , derive (38.3) and explain the multiplier.
For independent per-sample accuracy and , compute exact majority-vote accuracy and best-of- accuracy.
Implement temperature distillation, gradient-check it with finite differences, and implement exact best-of- and majority-vote accuracy using the binomial distribution.
References
-
[agarwal2023onpolicy] R. Agarwal et al. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. 2023. arXiv:2306.13649
-
[deepseekai2025deepseekr1] DeepSeek-AI et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. arXiv:2501.12948
-
[hinton2015distilling] G. Hinton, O. Vinyals, and J. Dean. Distilling the Knowledge in a Neural Network. 2015. arXiv:1503.02531
-
[lightman2023let] H. Lightman et al. Let’s Verify Step by Step. 2023. arXiv:2305.20050
-
[snell2024scaling] C. Snell et al. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. 2024. arXiv:2408.03314
-
[wang2022selfconsistency] X. Wang et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. arXiv:2203.11171