Chapter 9

Learning from Data

Linear and logistic regression, MSE and cross-entropy from maximum likelihood, and gradient descent.

Learning from data means choosing parameters that make a model behave well on examples it has not memorized. For LLMs in 2026, the examples might be next-token targets, instruction responses, preference labels, or benchmark tasks converted into losses. This chapter uses tiny supervised problems to show the core loop: define a loss, compute its gradient, optimize on training data, and use held-out data to decide whether the model learned the pattern.

9.1 Supervised learning and empirical risk

A supervised dataset is {(xi,yi)}i=1N\{(\vx_i, y_i)\}_{i=1}^N, with features xi\vx_i and target yiy_i. A model fθ(x)f_\vtheta(\vx) maps features to a prediction. A loss ℓ(fθ(x),y)\ell(f_\vtheta(\vx), y) says how bad one prediction is. The population risk is the expected loss on future data, but we only see samples, so training minimizes empirical risk:

R^(θ)=1N∑i=1Nℓ(fθ(xi),yi).(9.1)\hat R(\vtheta) = \frac1N \sum_{i=1}^N \ell(f_\vtheta(\vx_i), y_i) .\tag{9.1}

This is the same sample-mean logic as Section 8.1: the training set estimates future performance, with noise and bias from how it was collected. Optimization lowers the empirical risk; validation checks whether that also lowered risk on unseen examples. In a language model, the feature can be the context tokens and the label the next token; in instruction tuning, the feature is the prompt and the label is a desired response. The mathematics is unchanged: a row produces a prediction, a loss turns the prediction into a scalar, and training averages those scalars.

9.2 Linear regression

Linear regression predicts y^=x⊤w+b\hat y = \vx^\T\vw + b and usually uses mean squared error. Write X\mX for the design matrix and append a column of ones so the bias is included in θ\vtheta. The objective is

L(θ)=1N∥Xθ−y∥22.(9.2)L(\vtheta) = \frac1N\|\mX\vtheta - \vy\|_2^2 .\tag{9.2}

Differentiate by expanding the square: ∇θL=2NX⊤(Xθ−y)\nabla_\vtheta L = \frac2N \mX^\T(\mX\vtheta-\vy). Setting the gradient to zero gives the normal equations,

X⊤Xθ=X⊤y.(9.3)\mX^\T\mX\vtheta = \mX^\T\vy .\tag{9.3}

For small full-rank problems, solving (9.3) gives the exact least-squares solution. The closed form is also a useful check on implementations: if gradient descent cannot match it on a tiny problem, the gradient or learning rate is wrong. For large models, forming X⊤X\mX^\T\mX is too expensive and a closed form no longer exists, so we update parameters by gradient descent. The learning rate η\eta sets the step size: too small wastes compute, while too large bounces or diverges. The actual vector update is:

θ←θ−η∇θL.(9.4)\vtheta \leftarrow \vtheta - \eta\nabla_\vtheta L .\tag{9.4}
Listing 9.1 Linear regression, its gradient, and the normal equations
def predict_linear(X, w, b):
    return X @ w + b


def mse_loss(X, y, w, b):
    residual = predict_linear(X, w, b) - y
    return float(np.mean(residual ** 2))


def mse_gradients(X, y, w, b):
    residual = predict_linear(X, w, b) - y
    grad_w = 2.0 * X.T @ residual / len(X)
    grad_b = 2.0 * residual.mean()
    return grad_w, grad_b


def normal_equation(X, y):
    X1 = np.c_[X, np.ones(len(X), dtype=X.dtype)]
    theta = np.linalg.solve(X1.T @ X1, X1.T @ y)
    return theta[:-1], theta[-1]
Listing 9.2 Full-batch gradient descent
def gradient_descent_linear(X, y, w, b, steps=200, lr=0.1):
    w = np.array(w, dtype=np.float32, copy=True)
    b = np.array(b, dtype=np.float32)
    losses = []
    for _ in range(steps):
        grad_w, grad_b = mse_gradients(X, y, w, b)
        w -= lr * grad_w.astype(np.float32)
        b -= np.float32(lr * grad_b)
        losses.append(mse_loss(X, y, w, b))
    return w, float(b), np.array(losses, dtype=np.float32)

The tests finite-difference the MSE gradient, then show gradient descent reaches the same solution as the normal equations on a small float32 problem.

9.3 Logistic regression

Binary logistic regression predicts a probability. First compute a logit z=x⊤w+bz = \vx^\T\vw + b, then p=σ(z)=1/(1+e−z)p = \sigma(z) = 1/(1+e^{-z}). The binary cross-entropy loss is

ℓ(p,y)=−ylog⁡p−(1−y)log⁡(1−p).(9.5)\ell(p,y) = -y\log p - (1-y)\log(1-p) .\tag{9.5}

The useful cancellation is why this model is everywhere. Since ∂ℓ/∂z=p−y\partial \ell / \partial z = p-y, averaging over rows gives

wˉ=1NX⊤(p−y),bˉ=1N∑i(pi−yi).(9.6)\bar{\vw} = \frac1N \mX^\T(\vp-\vy), \qquad \bar b = \frac1N\sum_i(p_i-y_i) .\tag{9.6}
Listing 9.3 Logistic regression loss and gradient
def sigmoid(z):
    positive = z >= 0
    out = np.empty_like(z, dtype=np.float64)
    out[positive] = 1.0 / (1.0 + np.exp(-z[positive]))
    exp_z = np.exp(z[~positive])
    out[~positive] = exp_z / (1.0 + exp_z)
    return out


def logistic_loss_and_grads(X, y, w, b):
    logits = X @ w + b
    p = sigmoid(logits)
    eps = 1e-12
    loss = -np.mean(y * np.log(p + eps) + (1.0 - y) * np.log(1.0 - p + eps))
    grad_logits = (p - y) / len(X)
    grad_w = X.T @ grad_logits
    grad_b = grad_logits.sum()
    return float(loss), grad_w, float(grad_b)

The chapter test checks exactly X⊤(p−y)/N\mX^\T(\vp-\vy)/N against finite differences. The sign is intuitive: when the model predicts too large a probability, p−yp-y is positive and the update lowers logits for features present in that row. This is the same pattern that becomes softmax cross-entropy for multi-class token prediction.

9.4 Optimization and data splits

Full-batch gradient descent computes the exact training gradient each step. Minibatch SGD samples a small batch, computes a noisy gradient, and updates immediately. The noise makes the path jagged, but each update is cheap and the average direction is right when batches are sampled fairly. Shuffling matters: batches sorted by topic, length, or difficulty can create biased updates that look like optimization bugs. This is why deep learning trains with minibatches rather than one exact pass over all tokens before every update, following the stochastic approximation idea of Robbins and Monro [robbins1951].

Listing 9.4 Minibatch SGD for logistic regression
def fit_logistic_minibatch(X, y, w, b, steps, lr, batch_size, rng):
    w = np.array(w, dtype=np.float32, copy=True)
    b = np.array(b, dtype=np.float32)
    losses = []
    for _ in range(steps):
        idx = rng.integers(0, len(X), size=batch_size)
        _, grad_w, grad_b = logistic_loss_and_grads(X[idx], y[idx], w, b)
        w -= lr * grad_w.astype(np.float32)
        b -= np.float32(lr * grad_b)
        losses.append(logistic_loss_and_grads(X, y, w, b)[0])
    return w, float(b), np.array(losses, dtype=np.float32)

Use three disjoint splits. The training set fits parameters. The validation set chooses model class, degree, learning rate, stopping time, and regularization. The test set is touched once, after those choices are frozen. Early stopping is just another validation choice: stop when validation loss stops improving, not when training loss is smallest. If the data distribution changes, make a new split that reflects the new deployment population.

Listing 9.5 A deterministic train/validation/test split
def train_validation_test_split(n, train=0.6, validation=0.2, rng=None):
    rng = np.random.default_rng(0) if rng is None else rng
    order = rng.permutation(n)
    n_train = int(round(train * n))
    n_validation = int(round(validation * n))
    train_idx = order[:n_train]
    validation_idx = order[n_train:n_train + n_validation]
    test_idx = order[n_train + n_validation:]
    return train_idx, validation_idx, test_idx

9.5 Underfitting, overfitting, and L2

Capacity is the set of functions a model can represent. A line underfits a curved relation: training and validation losses are both high because the model class is too small. A very high degree polynomial can interpolate noisy training targets and overfit: training loss is low, but validation loss rises because the wiggles are noise. A middle degree can match the signal. This example is deliberately small, but the same shape appears in neural networks: too little capacity misses structure, too much capacity can fit spurious details unless data, regularization, or early stopping constrain it.

Listing 9.6 Polynomial features for capacity experiments
def polynomial_features(x, degree):
    x = np.asarray(x)
    powers = [x ** power for power in range(1, degree + 1)]
    return np.stack(powers, axis=1)

L2 regularization adds a weight penalty to empirical risk,

R^λ(θ)=R^(θ)+λ∥w∥22.(9.7)\hat R_\lambda(\vtheta) = \hat R(\vtheta) + \lambda\|\vw\|_2^2 .\tag{9.7}

The penalty discourages large weights, which damps the wiggles of high-capacity models and usually improves validation loss when data are scarce. The bias term is often left unpenalized. Regularization is not a substitute for a validation set; λ\lambda itself is a choice that must be selected without looking at the test set.

In practice

Modern LLM training still follows this supervised template at several scales: next-token pretraining minimizes cross-entropy over text, and instruction tuning minimizes supervised losses on prompt-response data, as in InstructGPT [ouyang2022training] and later Llama 3 training reports [grattafiori2024]. Production teams rarely trust training loss alone; they track held-out perplexity, task metrics, and safety or preference evaluations. Validation sets are also where early stopping and data-mixture choices are made. Once a benchmark steers those choices, it is no longer a clean test set.

Key equations
R^(θ)=1N∑iℓ(fθ(xi),yi)\hat R(\vtheta) = \frac1N\sum_i \ell(f_\vtheta(\vx_i), y_i)
∇θ1N∥Xθ−y∥2=2NX⊤(Xθ−y)\nabla_\vtheta \frac1N\|\mX\vtheta-\vy\|^2 = \frac2N\mX^\T(\mX\vtheta-\vy)
X⊤Xθ=X⊤y\mX^\T\mX\vtheta = \mX^\T\vy
θ←θ−η∇θL\vtheta \leftarrow \vtheta - \eta\nabla_\vtheta L
∇wLBCE=1NX⊤(p−y)\nabla_\vw L_{\mathrm{BCE}} = \frac1N\mX^\T(\vp-\vy)

9.6 Teach it

The one-sentence version: supervised learning fits parameters by minimizing average loss on training examples, then uses held-out examples to detect whether the pattern generalized. Analogy: training is practicing with answer keys; validation is a mock exam used to choose how to study; the test is the final exam you do not peek at. Board steps: write empirical risk; derive a linear-regression gradient; show the logistic BCE cancellation p−yp-y; draw underfit, good fit, and overfit polynomial curves. Misconceptions: low training loss is not the goal; SGD is not a different objective from gradient descent; the test set is not for tuning. Check for understanding: why does the gradient for logistic regression contain the feature matrix transpose?

9.7 Exercises

Exercise 9.1 ★ Empirical risk

Explain why empirical risk is an estimate of population risk. Name one way this estimate can be biased even with many examples.

Exercise 9.2 ★★ Normal equations

Derive X⊤Xθ=X⊤y\mX^\T\mX\vtheta = \mX^\T\vy from the mean squared error objective with a bias column already appended to X\mX.

Exercise 9.3 ★★ BCE gradient

For binary logistic regression, derive ∇wL=X⊤(p−y)/N\nabla_\vw L = \mX^\T(\vp-\vy)/N. Why is the result a vector with the same shape as w\vw?

Exercise 9.4 ★★★ Implementation

Fit polynomial regressions of degrees 1, 3, and 11 to a tiny noisy training set and compare validation losses. Which one underfits, which one overfits, and which one best matches the signal?

References

  • [goodfellow2016] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. https://www.deeplearningbook.org

  • [bishop2006] C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.

  • [grattafiori2024] A. Grattafiori et al. The Llama 3 herd of models. 2024. arXiv:2407.21783

  • [robbins1951] H. Robbins and S. Monro. A stochastic approximation method. Annals of Mathematical Statistics 22(3), 400–407, 1951.

  • [ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155