Chapter 13
Loss Functions & Divergences
MSE, Huber, cross-entropy, forward and reverse KL, focal loss, and knowledge distillation.
A loss is the scalar story a model tells the optimizer. In supervised learning it is usually a negative log-likelihood: choose a noise model for the target, take minus the log probability of the observation, and differentiate. This chapter connects that view to regression losses, binary classification, KL-style distribution matching, and knowledge distillation. Contrastive losses are left to Chapter 30.
13.1 Losses as negative log-likelihoods
Maximum likelihood chooses parameters that make the observed data likely (Section 6.5). Minimizing a loss is the same procedure after dropping constants that do not depend on the prediction. If with fixed , then
The prediction-dependent term is squared error. If the noise is Laplace, , the loss is absolute error plus a constant. The probability model is not decoration: it says what kind of residuals the model expects and how harshly it treats outliers.
For a dataset, the objective is the mean of these per-example negative log-likelihoods. A constant can be dropped for optimization, but the scale still matters when losses are combined: doubling a loss doubles its gradient. That is why "MSE" in code must be read with its exact normalization. Some libraries use , others use , and their optima match but their learning-rate needs differ.
13.2 MSE, MAE, and Huber
Let . Mean squared error, mean absolute error, and Huber loss are
Their gradients with respect to the prediction are , away from zero, and inside the Huber quadratic region but outside it [huber1964]. MSE keeps increasing the gradient as an outlier moves farther away, so one bad example can dominate a small batch. MAE caps every nonzero residual at the same gradient size, which is robust but has a kink at zero. Huber is the compromise: quadratic near the optimum, linear in the tails.
This is a robustness statement about gradients, not only about loss values. For residual , MSE pushes with gradient 20, while MAE and Huber with push with gradient 1. If the large residual is a mislabeled example, the capped gradient protects the rest of the batch. If it is a genuine rare case, the cap slows learning on exactly the example you may care about. The loss encodes that trade-off.
def mse_loss(prediction, target):
residual = np.asarray(prediction, dtype=np.float64) - target
return _mean_loss_and_grad(residual ** 2, 2 * residual)
def mae_loss(prediction, target):
residual = np.asarray(prediction, dtype=np.float64) - target
return _mean_loss_and_grad(np.abs(residual), np.sign(residual))
def huber_loss(prediction, target, delta=1.0):
residual = np.asarray(prediction, dtype=np.float64) - target
abs_r = np.abs(residual)
quadratic = abs_r <= delta
loss = np.where(quadratic, 0.5 * residual ** 2,
delta * (abs_r - 0.5 * delta))
grad = np.where(quadratic, residual, delta * np.sign(residual))
return _mean_loss_and_grad(loss, grad)
13.3 Binary cross-entropy and focal loss
For a binary label and logit , binary cross-entropy is the negative log-likelihood of a Bernoulli with probability :
Using algebra and the same stability idea as softmax, this becomes
The stable form never computes after has rounded to 1. Focal loss adds a factor that downweights easy examples [lin2017focal]. With for and for ,
When is already near 1, the multiplier is tiny; when the example is misclassified, the loss behaves much more like cross-entropy.
The parameter controls how aggressively easy examples are suppressed; setting recovers weighted BCE. The optional balances positive and negative classes. Focal loss is therefore not a generic "better BCE." It is a targeted fix for class imbalance where the training signal would otherwise be flooded by many already-correct examples.
def sigmoid(x):
x = np.asarray(x, dtype=np.float64)
z = np.exp(-np.abs(x))
return np.where(x >= 0, 1 / (1 + z), z / (1 + z))
def bce_with_logits(logits, targets):
logits = np.asarray(logits, dtype=np.float64)
targets = np.asarray(targets, dtype=np.float64)
loss = np.maximum(logits, 0) - logits * targets
loss = loss + np.log1p(np.exp(-np.abs(logits)))
return _mean_loss_and_grad(loss, sigmoid(logits) - targets)
def binary_focal_loss(logits, targets, gamma=2.0, alpha=0.25):
logits = np.asarray(logits, dtype=np.float64)
targets = np.asarray(targets, dtype=np.float64)
sign = 2 * targets - 1
log_pt = -np.logaddexp(0, -sign * logits)
pt = np.exp(log_pt)
alpha_t = alpha * targets + (1 - alpha) * (1 - targets)
loss = -alpha_t * (1 - pt) ** gamma * log_pt
dloss_dpt = alpha_t * gamma * (1 - pt) ** (gamma - 1) * log_pt
dloss_dpt = dloss_dpt - alpha_t * (1 - pt) ** gamma / pt
grad = dloss_dpt * sign * pt * (1 - pt)
return _mean_loss_and_grad(loss, grad)
13.4 Divergences as losses
Cross-entropy differs from forward KL by the target entropy, which is constant when the target distribution is fixed (Chapter 7). Thus minimizing over model is the same as fitting the target by cross-entropy, and its logit gradient is . Reverse KL, , averages over the model’s own distribution. Its gradient depends on where the model already puts mass, so it is more mode-seeking and can ignore target modes it does not sample.
As losses, the two directions answer different questions. Forward KL asks the model to cover everything the target assigns probability to; putting near zero where is positive is expensive. Reverse KL asks whether the model’s own samples look plausible under ; it is less bothered by target regions the model never visits. This distinction is why maximum-likelihood training, distillation, variational inference, and policy regularization can all say "KL" while behaving differently.
Jensen-Shannon divergence symmetrizes KL by comparing each distribution with their midpoint, :
It is finite and symmetric, but in this book it mostly appears as a diagnostic; training losses usually use cross-entropy or a directed KL.
The midpoint also prevents the infinite value that ordinary KL gets when one distribution has support where the other has zero. That makes Jensen-Shannon easier to plot and compare, but its symmetry removes the useful modelling choice of deciding which distribution supplies the expectation.
def kl_forward_logits(target_probs, logits):
p = np.asarray(target_probs, dtype=np.float64)
log_q = log_softmax(logits)
loss = np.sum(p * (np.log(p) - log_q), axis=-1)
return float(np.mean(loss)), (np.exp(log_q) - p) / logits.shape[0]
def kl_reverse_logits(logits, target_probs):
q = softmax(logits)
log_q = log_softmax(logits)
log_p = np.log(np.asarray(target_probs, dtype=np.float64))
values = log_q - log_p + 1
loss = np.sum(q * (log_q - log_p), axis=-1)
centered = values - np.sum(q * values, axis=-1, keepdims=True)
return float(np.mean(loss)), q * centered / logits.shape[0]
def jensen_shannon(p, q):
p, q = np.asarray(p, dtype=np.float64), np.asarray(q, dtype=np.float64)
m = 0.5 * (p + q)
return 0.5 * np.sum(p * (np.log(p) - np.log(m))) + 0.5 * np.sum(
q * (np.log(q) - np.log(m))
)
13.5 Knowledge distillation
Knowledge distillation trains a student to match a teacher distribution rather than only the hard label [hinton2015distilling]. Let and . The usual loss is a temperature-scaled cross-entropy or forward KL:
Without the , the gradient with respect to the student logits would be . For large , both softened distributions move toward uniform and their difference is , so the unscaled gradient is . Multiplying the loss by gives the implemented gradient
up to the batch mean. The scale keeps the distillation signal comparable as changes.
Soft targets carry information that a one-hot label discards. If the teacher assigns a little probability to several similar classes, the student sees that structure in every update. The temperature makes those dark probabilities visible by flattening the teacher distribution. The factor then prevents the visible signal from shrinking just because the softening temperature was increased.
def distillation_loss(student_logits, teacher_logits, temperature=2.0, scale=True):
teacher = softmax(teacher_logits, temperature=temperature)
student_log_p = log_softmax(student_logits, temperature=temperature)
batch = student_logits.shape[0]
factor = temperature ** 2 if scale else 1.0
loss = -factor * np.sum(teacher * student_log_p) / batch
student = np.exp(student_log_p)
grad = factor * (student - teacher) / (batch * temperature)
return float(loss), grad
|
In practice
|
Pretraining and supervised fine-tuning of LLMs are maximum-likelihood training with softmax cross-entropy. Regression heads choose MSE, MAE, or Huber according to the assumed residual noise and desired outlier robustness. Focal loss is mainly used when easy negatives overwhelm rare positives, as in dense detection [lin2017focal]. Distillation uses a teacher distribution and often a temperature, so it can transfer relative preferences among wrong classes, not just the top label [hinton2015distilling]. |
13.6 Teach it
The one-sentence version. A loss is usually a negative log-likelihood; its gradient says how the prediction should move to make the observed target less surprising.
An analogy. Choosing a loss is choosing a judge. MSE is a judge who shouts louder as an error gets larger; MAE speaks at the same volume for every miss; Huber shouts near the target and then caps its voice.
At the board.
-
Write Gaussian NLL and cross out constants to reveal squared error.
-
Write Laplace NLL and reveal absolute error.
-
Plot MSE, MAE, and Huber gradients against one large residual.
-
For distillation, write , then explain why is added.
Misconceptions to address.
-
"Loss names are arbitrary." Most encode a probability model or divergence direction.
-
"Robust means ignoring errors." Robust losses still move outliers, but cap their leverage.
-
"Forward and reverse KL are interchangeable." Their averaging distributions differ.
Check for understanding. Which loss would you choose if 1% of labels are huge measurement errors, and why?
13.7 Exercises
Starting from Gaussian and Laplace likelihoods with fixed scale, show why MSE and MAE are negative log-likelihood losses up to constants.
Derive the gradients of MSE, MAE, and Huber loss with respect to the prediction. Then derive the stable BCE-with-logits form and its gradient .
Explain how focal loss changes binary cross-entropy when an example is already classified correctly. Then compare forward KL and reverse KL as losses, and state why Jensen-Shannon is symmetric.
Derive the gradient of temperature-distillation loss with and without the multiplier.
Use distillation_loss to gradient-check a tiny student logit matrix and verify the scaling.
References
-
[goodfellow2016] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. https://www.deeplearningbook.org
-
[bishop2006] C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
-
[hinton2015distilling] G. Hinton, O. Vinyals, and J. Dean. Distilling the Knowledge in a Neural Network. 2015. arXiv:1503.02531
-
[huber1964] P. J. Huber. Robust estimation of a location parameter. Annals of Mathematical Statistics 35(1), 73-101, 1964.
-
[lin2017focal] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár. Focal Loss for Dense Object Detection. 2017. arXiv:1708.02002