Chapter 9
Learning from Data
Linear and logistic regression, MSE and cross-entropy from maximum likelihood, and gradient descent.
Learning from data means choosing parameters that make a model behave well on examples it has not memorized. For LLMs in 2026, the examples might be next-token targets, instruction responses, preference labels, or benchmark tasks converted into losses. This chapter uses tiny supervised problems to show the core loop: define a loss, compute its gradient, optimize on training data, and use held-out data to decide whether the model learned the pattern.
9.1 Supervised learning and empirical risk
A supervised dataset is , with features and target . A model maps features to a prediction. A loss says how bad one prediction is. The population risk is the expected loss on future data, but we only see samples, so training minimizes empirical risk:
This is the same sample-mean logic as Section 8.1: the training set estimates future performance, with noise and bias from how it was collected. Optimization lowers the empirical risk; validation checks whether that also lowered risk on unseen examples. In a language model, the feature can be the context tokens and the label the next token; in instruction tuning, the feature is the prompt and the label is a desired response. The mathematics is unchanged: a row produces a prediction, a loss turns the prediction into a scalar, and training averages those scalars.
9.2 Linear regression
Linear regression predicts and usually uses mean squared error. Write for the design matrix and append a column of ones so the bias is included in . The objective is
Differentiate by expanding the square: . Setting the gradient to zero gives the normal equations,
For small full-rank problems, solving (9.3) gives the exact least-squares solution. The closed form is also a useful check on implementations: if gradient descent cannot match it on a tiny problem, the gradient or learning rate is wrong. For large models, forming is too expensive and a closed form no longer exists, so we update parameters by gradient descent. The learning rate sets the step size: too small wastes compute, while too large bounces or diverges. The actual vector update is:
def predict_linear(X, w, b):
return X @ w + b
def mse_loss(X, y, w, b):
residual = predict_linear(X, w, b) - y
return float(np.mean(residual ** 2))
def mse_gradients(X, y, w, b):
residual = predict_linear(X, w, b) - y
grad_w = 2.0 * X.T @ residual / len(X)
grad_b = 2.0 * residual.mean()
return grad_w, grad_b
def normal_equation(X, y):
X1 = np.c_[X, np.ones(len(X), dtype=X.dtype)]
theta = np.linalg.solve(X1.T @ X1, X1.T @ y)
return theta[:-1], theta[-1]
def gradient_descent_linear(X, y, w, b, steps=200, lr=0.1):
w = np.array(w, dtype=np.float32, copy=True)
b = np.array(b, dtype=np.float32)
losses = []
for _ in range(steps):
grad_w, grad_b = mse_gradients(X, y, w, b)
w -= lr * grad_w.astype(np.float32)
b -= np.float32(lr * grad_b)
losses.append(mse_loss(X, y, w, b))
return w, float(b), np.array(losses, dtype=np.float32)
The tests finite-difference the MSE gradient, then show gradient descent reaches the same solution as the normal equations on a small float32 problem.
9.3 Logistic regression
Binary logistic regression predicts a probability. First compute a logit , then . The binary cross-entropy loss is
The useful cancellation is why this model is everywhere. Since , averaging over rows gives
def sigmoid(z):
positive = z >= 0
out = np.empty_like(z, dtype=np.float64)
out[positive] = 1.0 / (1.0 + np.exp(-z[positive]))
exp_z = np.exp(z[~positive])
out[~positive] = exp_z / (1.0 + exp_z)
return out
def logistic_loss_and_grads(X, y, w, b):
logits = X @ w + b
p = sigmoid(logits)
eps = 1e-12
loss = -np.mean(y * np.log(p + eps) + (1.0 - y) * np.log(1.0 - p + eps))
grad_logits = (p - y) / len(X)
grad_w = X.T @ grad_logits
grad_b = grad_logits.sum()
return float(loss), grad_w, float(grad_b)
The chapter test checks exactly against finite differences. The sign is intuitive: when the model predicts too large a probability, is positive and the update lowers logits for features present in that row. This is the same pattern that becomes softmax cross-entropy for multi-class token prediction.
9.4 Optimization and data splits
Full-batch gradient descent computes the exact training gradient each step. Minibatch SGD samples a small batch, computes a noisy gradient, and updates immediately. The noise makes the path jagged, but each update is cheap and the average direction is right when batches are sampled fairly. Shuffling matters: batches sorted by topic, length, or difficulty can create biased updates that look like optimization bugs. This is why deep learning trains with minibatches rather than one exact pass over all tokens before every update, following the stochastic approximation idea of Robbins and Monro [robbins1951].
def fit_logistic_minibatch(X, y, w, b, steps, lr, batch_size, rng):
w = np.array(w, dtype=np.float32, copy=True)
b = np.array(b, dtype=np.float32)
losses = []
for _ in range(steps):
idx = rng.integers(0, len(X), size=batch_size)
_, grad_w, grad_b = logistic_loss_and_grads(X[idx], y[idx], w, b)
w -= lr * grad_w.astype(np.float32)
b -= np.float32(lr * grad_b)
losses.append(logistic_loss_and_grads(X, y, w, b)[0])
return w, float(b), np.array(losses, dtype=np.float32)
Use three disjoint splits. The training set fits parameters. The validation set chooses model class, degree, learning rate, stopping time, and regularization. The test set is touched once, after those choices are frozen. Early stopping is just another validation choice: stop when validation loss stops improving, not when training loss is smallest. If the data distribution changes, make a new split that reflects the new deployment population.
def train_validation_test_split(n, train=0.6, validation=0.2, rng=None):
rng = np.random.default_rng(0) if rng is None else rng
order = rng.permutation(n)
n_train = int(round(train * n))
n_validation = int(round(validation * n))
train_idx = order[:n_train]
validation_idx = order[n_train:n_train + n_validation]
test_idx = order[n_train + n_validation:]
return train_idx, validation_idx, test_idx
9.5 Underfitting, overfitting, and L2
Capacity is the set of functions a model can represent. A line underfits a curved relation: training and validation losses are both high because the model class is too small. A very high degree polynomial can interpolate noisy training targets and overfit: training loss is low, but validation loss rises because the wiggles are noise. A middle degree can match the signal. This example is deliberately small, but the same shape appears in neural networks: too little capacity misses structure, too much capacity can fit spurious details unless data, regularization, or early stopping constrain it.
def polynomial_features(x, degree):
x = np.asarray(x)
powers = [x ** power for power in range(1, degree + 1)]
return np.stack(powers, axis=1)
L2 regularization adds a weight penalty to empirical risk,
The penalty discourages large weights, which damps the wiggles of high-capacity models and usually improves validation loss when data are scarce. The bias term is often left unpenalized. Regularization is not a substitute for a validation set; itself is a choice that must be selected without looking at the test set.
|
In practice
|
Modern LLM training still follows this supervised template at several scales: next-token pretraining minimizes cross-entropy over text, and instruction tuning minimizes supervised losses on prompt-response data, as in InstructGPT [ouyang2022training] and later Llama 3 training reports [grattafiori2024]. Production teams rarely trust training loss alone; they track held-out perplexity, task metrics, and safety or preference evaluations. Validation sets are also where early stopping and data-mixture choices are made. Once a benchmark steers those choices, it is no longer a clean test set. |
9.6 Teach it
The one-sentence version: supervised learning fits parameters by minimizing average loss on training examples, then uses held-out examples to detect whether the pattern generalized. Analogy: training is practicing with answer keys; validation is a mock exam used to choose how to study; the test is the final exam you do not peek at. Board steps: write empirical risk; derive a linear-regression gradient; show the logistic BCE cancellation ; draw underfit, good fit, and overfit polynomial curves. Misconceptions: low training loss is not the goal; SGD is not a different objective from gradient descent; the test set is not for tuning. Check for understanding: why does the gradient for logistic regression contain the feature matrix transpose?
9.7 Exercises
Explain why empirical risk is an estimate of population risk. Name one way this estimate can be biased even with many examples.
Derive from the mean squared error objective with a bias column already appended to .
For binary logistic regression, derive . Why is the result a vector with the same shape as ?
Fit polynomial regressions of degrees 1, 3, and 11 to a tiny noisy training set and compare validation losses. Which one underfits, which one overfits, and which one best matches the signal?
References
-
[goodfellow2016] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. https://www.deeplearningbook.org
-
[bishop2006] C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
-
[grattafiori2024] A. Grattafiori et al. The Llama 3 herd of models. 2024. arXiv:2407.21783
-
[robbins1951] H. Robbins and S. Monro. A stochastic approximation method. Annals of Mathematical Statistics 22(3), 400–407, 1951.
-
[ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155