Appendix A

Notation & Shapes

Symbols, typography, and the shape conventions used throughout the book.

This book writes mathematics the way the code is written. Examples are rows, the batch axis comes first, and a gradient has the shape of the thing it differentiates. This appendix fixes those conventions once, so the chapters can use them without comment.

A.1 How the chapters work

Understanding a model means passing three tests. Can you write it: state the equations and derive the gradients on a blank page? Can you code it: implement it in NumPy and verify it against a finite difference? Can you teach it: explain it so that someone else can pass the first two tests? Each chapter is built around those tests, in the same order:

  1. Why it matters: where the idea appears in current models.

  2. Intuition: a picture or analogy to hang the mathematics on.

  3. The math: definitions, the forward computation with shapes, then the backward pass.

  4. Code: tested NumPy listings, each checked against finite differences.

  5. In practice: how production models use or vary the idea, with primary sources.

  6. Key equations: a boxed summary, collected in Appendix E for review.

  7. Teach it: a one-sentence summary, an analogy, a board plan, common misconceptions, and questions to ask a learner.

  8. Exercises: ★ concepts, ★★ derivations, and ★★★ implementations. Worked solutions are in Appendix D.

A good way to study is to read with a pencil, then close the book and rederive the boxed equations. Next, type the listings rather than copying them, and run the gradient check. Only then open the solutions. Finally, give the "Teach it" explanation aloud, to a person or to an empty room.

A.2 Typography

Table A.1 Symbols used throughout the book
Symbol Meaning

a,x,ηa, x, \eta

Scalars: italic lowercase letters.

x,h,θ\vx, \vh, \vtheta

Vectors: bold lowercase letters.

X,W\mX, \mW

Matrices and higher-order arrays: bold uppercase letters.

xi,  Wijx_i, \; W_{ij}

Entries: italic, with indices. Mathematics counts from 1; code counts from 0.

Xi,:,  X:,j\mX_{i,:}, \; \mX_{:,j}

Row ii and column jj of X\mX, as in NumPy’s X[i, :] and X[:, j].

Rm×n\R^{m \times n}

Real arrays with mm rows and nn columns.

A⊤\mA^\T

Transpose.

a⋅b,  ⟨A,B⟩\va \cdot \vb, \; \langle \mA, \mB \rangle

Dot product; the inner product ∑ijAijBij\sum_{ij} A_{ij} B_{ij} of same-shaped arrays.

AB,  A⊙B\mA \mB, \; \mA \odot \mB

Matrix product (NumPy A @ B); elementwise product (A * B).

∥x∥\lVert \vx \rVert

Euclidean norm, x⋅x\sqrt{\vx \cdot \vx}.

1,  diag⁡(v)\one, \; \diag(\vv)

A vector of ones; the diagonal matrix with v\vv on its diagonal.

log⁡\log

Natural logarithm. log⁡2\log_2 appears when measuring in bits.

A.3 Shapes and the row convention

A batch of NN examples with dd features is a matrix X∈RN×d\mX \in \R^{N \times d} whose rows are examples. A linear layer maps each row with a weight matrix W∈Rdin×dout\mW \in \R^{d_\text{in} \times d_\text{out}} and bias b∈Rdout\vb \in \R^{d_\text{out}}:

Y=XW+b,Y∈RN×dout.(A.1)\mY = \mX \mW + \vb, \qquad \mY \in \R^{N \times d_\text{out}} .\tag{A.1}

This is exactly Y = X @ W + b. The bias broadcasts over rows. Many textbooks use the column convention y=Wx+b\vy = \mW \vx + \vb instead, with W\mW of shape dout×dind_\text{out} \times d_\text{in}. The two are transposes of each other. PyTorch’s nn.Linear stores its weight in the column-convention shape but computes on row batches, as x @ weight.T + bias [pytorch-linear].

The same letters name the same axes in every chapter:

Table A.2 Dimension names
Letter Axis

NN

Examples in a batch of independent examples

BB

Sequences in a batch of sequences

TT

Positions (tokens) in a sequence

dd

Model width: the size of each token’s hidden vector

dffd_\text{ff}

Hidden width of a feed-forward block

H,  dhH, \; d_h

Attention heads, and the width of each head (usually dh=d/Hd_h = d / H)

HkvH_\text{kv}

Key-value heads in grouped-query attention

VV

Vocabulary size

CC

Number of classes

LL

Number of layers

E,  kE, \; k

Experts in a mixture-of-experts layer, and experts chosen per token

Code comments annotate shapes in the same letters, as in # (B, T, d). When a listing’s shapes are unclear, the comments are the specification.

A.4 Derivatives and gradients

Training minimizes a scalar loss LL. For any array A\mA that LL depends on, the gradient of LL with respect to A\mA is written with a bar:

Aˉ=∂L∂A,Aˉij=∂L∂Aij.(A.2)\bar{\mA} = \frac{\partial L}{\partial \mA}, \qquad \bar{A}_{ij} = \frac{\partial L}{\partial A_{ij}} .\tag{A.2}

Aˉ\bar{\mA} has the same shape as A\mA. In code it is grad_A. During backpropagation, the gradient flowing into a layer from above is its upstream gradient.

For a function y=f(x)\vy = f(\vx) from Rn\R^n to Rm\R^m, the Jacobian J∈Rm×n\mJ \in \R^{m \times n} has entries Jij=∂yi/∂xjJ_{ij} = \partial y_i / \partial x_j: one row per output, one column per input. The chain rule then says:

xˉj=∑i∂L∂yi∂yi∂xj,that is,xˉ=J⊤yˉ.(A.3)\bar{x}_j = \sum_i \frac{\partial L}{\partial y_i} \frac{\partial y_i}{\partial x_j}, \qquad\text{that is,}\qquad \bar{\vx} = \mJ^\T \bar{\vy} .\tag{A.3}

This is a vector–Jacobian product. Backpropagation computes these products directly and almost never builds a Jacobian (Exercise A.4).

The affine layer is the worked example. Writing out one entry, Yik=∑jXijWjk+bkY_{ik} = \sum_j X_{ij} W_{jk} + b_k, and applying (A.3) entry by entry gives

Xˉ=YˉW⊤,Wˉ=X⊤Yˉ,bˉ=∑i=1NYˉi,:.(A.4)\bar{\mX} = \bar{\mY} \mW^\T, \qquad \bar{\mW} = \mX^\T \bar{\mY}, \qquad \bar{\vb} = \sum_{i=1}^{N} \bar{\mY}_{i,:} .\tag{A.4}

There is a quick way to remember these, but it only checks shapes, not correctness. Each gradient must have the shape of its variable, and Yˉ\bar{\mY} must appear exactly once. Wˉ\bar{\mW} has shape din×doutd_\text{in} \times d_\text{out}, and the only product of X\mX and Yˉ\bar{\mY} with that shape is X⊤Yˉ\mX^\T \bar{\mY}. The bias gradient is a sum because the bias was broadcast over rows (Section B.3).

Listing A.1 The affine layer, forward and backward
def affine_forward(X, W, b):
    """X (N, d_in), W (d_in, d_out), b (d_out,) -> Y (N, d_out)."""
    return X @ W + b


def affine_backward(grad_Y, X, W):
    """Given grad_Y = dL/dY (N, d_out), return dL/dX, dL/dW, dL/db."""
    grad_X = grad_Y @ W.T          # (N, d_out) @ (d_out, d_in) -> (N, d_in)
    grad_W = X.T @ grad_Y          # (d_in, N) @ (N, d_out)     -> (d_in, d_out)
    grad_b = grad_Y.sum(axis=0)    # b was broadcast over the N rows
    return grad_X, grad_W, grad_b

A.5 Probability and information

Table A.3 Probability notation
Symbol Meaning

p(x)p(x)

A probability (discrete xx) or density (continuous xx).

pθ(y∣x)p_\vtheta(y \mid x)

A model’s distribution over yy given xx, with parameters θ\vtheta.

x∼px \sim p

xx is drawn from pp.

Ex∼p[f(x)]\E_{x \sim p}[f(x)]

Expectation of f(x)f(x) when x∼px \sim p.

Var⁡[x],  Cov⁡[x,y]\Var[x], \; \Cov[x, y]

Variance and covariance.

H(p),  H(p,q)H(p), \; H(p, q)

Entropy of pp; cross-entropy of qq relative to pp.

DKL(p ∥ q)\KL(p \,\Vert\, q)

Kullback–Leibler divergence from qq to pp.

σ(x)\sigma(x)

The logistic sigmoid, 1/(1+e−x)1/(1+e^{-x}). In a probability context, a standard deviation.

softmax⁡(z)\softmax(\vz)

ezi/∑jezje^{z_i} / \sum_j e^{z_j}, applied along the last axis.

Predictions carry a hat, y^\hat{\vy}; targets do not. θ\vtheta collects all trainable parameters, η\eta is a learning rate, and tt counts optimization steps. Information is measured in nats, the unit of the natural logarithm, unless a chapter says bits.

A.6 Code conventions

  • Arrays are float32 unless stated otherwise. Gradient checks use float64 copies (Section B.8).

  • Randomness comes from one np.random.default_rng(seed) generator, passed explicitly.

  • A gradient variable is named after its variable, like grad_W for Wˉ\bar{\mW}.

  • Shared building blocks live in the scratch package. Each of its modules states the chapter that introduces it, and imports only from earlier chapters and the appendices.

Key equations
Aˉ=∂L∂A has the shape of A,xˉ=J⊤yˉ\bar{\mA} = \frac{\partial L}{\partial \mA} \text{ has the shape of } \mA, \qquad \bar{\vx} = \mJ^\T \bar{\vy}
Y=XW+b  ⟹  Xˉ=YˉW⊤,Wˉ=X⊤Yˉ,bˉ=∑iYˉi,:\mY = \mX \mW + \vb \;\Longrightarrow\; \bar{\mX} = \bar{\mY} \mW^\T,\quad \bar{\mW} = \mX^\T \bar{\mY},\quad \bar{\vb} = \textstyle\sum_i \bar{\mY}_{i,:}

A.7 Teach it

The one-sentence version. A gradient is a report card with one grade per number in the model. It has exactly the shape of the thing it grades.

An analogy. A loss is a single score for a whole team. The gradient tells each player, one entry per number, how much the team score would change if that number moved a little.

At the board.

  1. Write Y=XW+b\mY = \mX\mW + \vb and label every shape.

  2. Write one entry, Yik=∑jXijWjk+bkY_{ik} = \sum_j X_{ij} W_{jk} + b_k, and ask which entries of Y\mY a given WjkW_{jk} touches. The answer is a whole column, one entry per example.

  3. Sum those contributions to get Wˉjk=∑iXijYˉik\bar{W}_{jk} = \sum_i X_{ij} \bar{Y}_{ik}, then recognise the sum as a matrix product.

  4. Confirm it with the shape rule, and point out that the rule alone cannot tell X⊤Yˉ\mX^\T \bar{\mY} from a wrong formula of the same shape. That is why the finite difference check exists.

Misconceptions to address.

  • "The gradient of a matrix is a four-index Jacobian." It could be written that way, but for a scalar loss it collapses to one number per entry: an array shaped like the matrix.

  • "Row and column conventions give different networks." They give the same network with transposed weights.

Check for understanding. Why does the bias gradient involve a sum, while the weight gradient involves a matrix product?

A.8 Exercises

Exercise A.1 ★ Shapes of a small network

A batch X∈R32×128\mX \in \R^{32 \times 128} passes through H=max⁡(0,XW1+b1)\mH = \max(0, \mX\mW_1 + \vb_1) and then Z=HW2+b2\mZ = \mH\mW_2 + \vb_2, with W1∈R128×512\mW_1 \in \R^{128 \times 512} and W2∈R512×10\mW_2 \in \R^{512 \times 10}. Give the shapes of b1\vb_1, b2\vb_2, H\mH, Z\mZ, Wˉ1\bar{\mW}_1, and bˉ2\bar{\vb}_2, and count the trainable parameters.

Exercise A.2 ★ Reading PyTorch weights

PyTorch’s nn.Linear(d_in, d_out) stores weight with shape (d_out, d_in) and computes x @ weight.T + bias. How do you turn its parameters into this book’s W\mW and b\vb? What is the gradient of weight in terms of X\mX and Yˉ\bar{\mY}?

Exercise A.3 ★★ Deriving the input gradient

Derive Xˉ=YˉW⊤\bar{\mX} = \bar{\mY}\mW^\T from (A.3) by working with individual entries. Then check affine_backward against finite differences using the random-upstream trick from Section B.8.

Exercise A.4 ★★ A Jacobian you never build

For an elementwise function yi=f(xi)y_i = f(x_i) on Rn\R^n, write the Jacobian and the vector–Jacobian product. How much memory does building the Jacobian cost compared with computing the product directly?

References