Appendix A
Notation & Shapes
Symbols, typography, and the shape conventions used throughout the book.
This book writes mathematics the way the code is written. Examples are rows, the batch axis comes first, and a gradient has the shape of the thing it differentiates. This appendix fixes those conventions once, so the chapters can use them without comment.
A.1 How the chapters work
Understanding a model means passing three tests. Can you write it: state the equations and derive the gradients on a blank page? Can you code it: implement it in NumPy and verify it against a finite difference? Can you teach it: explain it so that someone else can pass the first two tests? Each chapter is built around those tests, in the same order:
-
Why it matters: where the idea appears in current models.
-
Intuition: a picture or analogy to hang the mathematics on.
-
The math: definitions, the forward computation with shapes, then the backward pass.
-
Code: tested NumPy listings, each checked against finite differences.
-
In practice: how production models use or vary the idea, with primary sources.
-
Key equations: a boxed summary, collected in Appendix E for review.
-
Teach it: a one-sentence summary, an analogy, a board plan, common misconceptions, and questions to ask a learner.
-
Exercises: ★ concepts, ★★ derivations, and ★★★ implementations. Worked solutions are in Appendix D.
A good way to study is to read with a pencil, then close the book and rederive the boxed equations. Next, type the listings rather than copying them, and run the gradient check. Only then open the solutions. Finally, give the "Teach it" explanation aloud, to a person or to an empty room.
A.2 Typography
| Symbol | Meaning |
|---|---|
Scalars: italic lowercase letters. |
|
Vectors: bold lowercase letters. |
|
Matrices and higher-order arrays: bold uppercase letters. |
|
Entries: italic, with indices. Mathematics counts from 1; code counts from 0. |
|
Row and column of , as in NumPy’s |
|
Real arrays with rows and columns. |
|
Transpose. |
|
Dot product; the inner product of same-shaped arrays. |
|
Matrix product (NumPy |
|
Euclidean norm, . |
|
A vector of ones; the diagonal matrix with on its diagonal. |
|
Natural logarithm. appears when measuring in bits. |
A.3 Shapes and the row convention
A batch of examples with features is a matrix whose rows are examples. A linear layer maps each row with a weight matrix and bias :
This is exactly Y = X @ W + b. The bias broadcasts over rows. Many textbooks use the column
convention instead, with of shape
. The two are transposes of each other. PyTorch’s
nn.Linear stores its weight in the column-convention shape but computes on row batches, as
x @ weight.T + bias [pytorch-linear].
The same letters name the same axes in every chapter:
| Letter | Axis |
|---|---|
Examples in a batch of independent examples |
|
Sequences in a batch of sequences |
|
Positions (tokens) in a sequence |
|
Model width: the size of each token’s hidden vector |
|
Hidden width of a feed-forward block |
|
Attention heads, and the width of each head (usually ) |
|
Key-value heads in grouped-query attention |
|
Vocabulary size |
|
Number of classes |
|
Number of layers |
|
Experts in a mixture-of-experts layer, and experts chosen per token |
Code comments annotate shapes in the same letters, as in # (B, T, d). When a listing’s
shapes are unclear, the comments are the specification.
A.4 Derivatives and gradients
Training minimizes a scalar loss . For any array that depends on, the gradient of with respect to is written with a bar:
has the same shape as . In code it is grad_A. During
backpropagation, the gradient flowing into a layer from above is its upstream gradient.
For a function from to , the Jacobian has entries : one row per output, one column per input. The chain rule then says:
This is a vector–Jacobian product. Backpropagation computes these products directly and almost never builds a Jacobian (Exercise A.4).
The affine layer is the worked example. Writing out one entry, , and applying (A.3) entry by entry gives
There is a quick way to remember these, but it only checks shapes, not correctness. Each gradient must have the shape of its variable, and must appear exactly once. has shape , and the only product of and with that shape is . The bias gradient is a sum because the bias was broadcast over rows (Section B.3).
def affine_forward(X, W, b):
"""X (N, d_in), W (d_in, d_out), b (d_out,) -> Y (N, d_out)."""
return X @ W + b
def affine_backward(grad_Y, X, W):
"""Given grad_Y = dL/dY (N, d_out), return dL/dX, dL/dW, dL/db."""
grad_X = grad_Y @ W.T # (N, d_out) @ (d_out, d_in) -> (N, d_in)
grad_W = X.T @ grad_Y # (d_in, N) @ (N, d_out) -> (d_in, d_out)
grad_b = grad_Y.sum(axis=0) # b was broadcast over the N rows
return grad_X, grad_W, grad_b
A.5 Probability and information
| Symbol | Meaning |
|---|---|
A probability (discrete ) or density (continuous ). |
|
A model’s distribution over given , with parameters . |
|
is drawn from . |
|
Expectation of when . |
|
Variance and covariance. |
|
Entropy of ; cross-entropy of relative to . |
|
Kullback–Leibler divergence from to . |
|
The logistic sigmoid, . In a probability context, a standard deviation. |
|
, applied along the last axis. |
Predictions carry a hat, ; targets do not. collects all trainable parameters, is a learning rate, and counts optimization steps. Information is measured in nats, the unit of the natural logarithm, unless a chapter says bits.
A.6 Code conventions
-
Arrays are
float32unless stated otherwise. Gradient checks usefloat64copies (Section B.8). -
Randomness comes from one
np.random.default_rng(seed)generator, passed explicitly. -
A gradient variable is named after its variable, like
grad_Wfor . -
Shared building blocks live in the
scratchpackage. Each of its modules states the chapter that introduces it, and imports only from earlier chapters and the appendices.
A.7 Teach it
The one-sentence version. A gradient is a report card with one grade per number in the model. It has exactly the shape of the thing it grades.
An analogy. A loss is a single score for a whole team. The gradient tells each player, one entry per number, how much the team score would change if that number moved a little.
At the board.
-
Write and label every shape.
-
Write one entry, , and ask which entries of a given touches. The answer is a whole column, one entry per example.
-
Sum those contributions to get , then recognise the sum as a matrix product.
-
Confirm it with the shape rule, and point out that the rule alone cannot tell from a wrong formula of the same shape. That is why the finite difference check exists.
Misconceptions to address.
-
"The gradient of a matrix is a four-index Jacobian." It could be written that way, but for a scalar loss it collapses to one number per entry: an array shaped like the matrix.
-
"Row and column conventions give different networks." They give the same network with transposed weights.
Check for understanding. Why does the bias gradient involve a sum, while the weight gradient involves a matrix product?
A.8 Exercises
A batch passes through and then , with and . Give the shapes of , , , , , and , and count the trainable parameters.
PyTorch’s nn.Linear(d_in, d_out) stores weight with shape (d_out, d_in) and computes
x @ weight.T + bias. How do you turn its parameters into this book’s and
? What is the gradient of weight in terms of and ?
Derive from (A.3) by working with individual entries.
Then check affine_backward against finite differences using the random-upstream trick from
Section B.8.
For an elementwise function on , write the Jacobian and the vector–Jacobian product. How much memory does building the Jacobian cost compared with computing the product directly?
References
-
[goodfellow2016] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. https://www.deeplearningbook.org
-
[parr2018] T. Parr and J. Howard. The matrix calculus you need for deep learning. 2018. arXiv:1802.01528
-
[pytorch-linear] PyTorch documentation.
torch.nn.Linear. https://docs.pytorch.org/docs/stable/generated/torch.nn.Linear.html