Appendix C
Matrix Calculus Cookbook
Vector-Jacobian products for the operations used in this book.
Matrix calculus is the bookkeeping layer behind backpropagation. This appendix is a compact reference for the NumPy operations used in the book: write the forward pass, receive an upstream cotangent, and return vector-Jacobian products (VJPs) with the input shapes. All formulas follow the gradient notation in Section A.4.
C.1 Method
Use index notation when unsure. If and the upstream gradient is , then
The shape rule is the guardrail: each returned VJP has the same shape as the primal it belongs to.
Broadcasted axes are summed away with unbroadcast. Reductions broadcast the upstream gradient back
to the input. VJPs compose backward through the graph; no full Jacobian is materialized.
def add_forward(x, y):
return x + y
def add_vjp(grad, x, y):
return unbroadcast(grad, x.shape), unbroadcast(grad, y.shape)
def multiply_forward(x, y):
return x * y
def multiply_vjp(grad, x, y):
return unbroadcast(grad * y, x.shape), unbroadcast(grad * x, y.shape)
def matmul_forward(a, b):
return a @ b
def matmul_vjp(grad, a, b):
return grad @ b.T, a.T @ grad
C.2 Cookbook
Let be the upstream gradient.
-
Add with broadcasting. . VJP: , .
-
Elementwise multiply. . VJP: , .
-
Matmul. . From , and .
-
Sum/mean. copies to every reduced element. Mean divides that copy by the number of reduced elements.
-
Exp/log/ReLU/sigmoid/tanh. VJPs multiply by the scalar derivative: , , , , and .
-
Softmax. For , .
-
Log-softmax and cross-entropy. If , . For mean cross-entropy, .
def softmax_forward(x, axis=-1):
shifted = x - np.max(x, axis=axis, keepdims=True)
exp = np.exp(shifted)
return exp / np.sum(exp, axis=axis, keepdims=True)
def softmax_vjp(grad, y, axis=-1):
dot = np.sum(grad * y, axis=axis, keepdims=True)
return y * (grad - dot)
-
LayerNorm. Over width , , . With , , where . Also , .
-
RMSNorm. . The VJP is like LayerNorm without centering; the code uses .
-
Embedding gather. . Scatter-add each upstream row into at its selected index; repeated IDs add.
-
Reshape/transpose. Reshape sends to the original shape. Transpose applies the inverse permutation.
-
Scaled dot-product attention. , , . VJPs: , , , , .
def scaled_dot_product_attention(q, k, v):
scale = 1 / np.sqrt(q.shape[-1])
scores = (q @ np.swapaxes(k, -1, -2)) * scale
probabilities = softmax_forward(scores, axis=-1)
return probabilities @ v
def attention_vjp(grad, q, k, v):
scale = 1 / np.sqrt(q.shape[-1])
scores = (q @ np.swapaxes(k, -1, -2)) * scale
probabilities = softmax_forward(scores, axis=-1)
grad_v = np.swapaxes(probabilities, -1, -2) @ grad
grad_prob = grad @ np.swapaxes(v, -1, -2)
grad_scores = softmax_vjp(grad_prob, probabilities, axis=-1)
grad_q = (grad_scores @ k) * scale
grad_k = (np.swapaxes(grad_scores, -1, -2) @ q) * scale
return grad_q, grad_k, grad_v
|
In practice
|
Autodiff systems implement these VJPs as kernels or kernel graphs, then gradient-check tricky new ops with finite differences [baydin2015automatic]. Transformers depend especially on LayerNorm [ba2016layer], RMSNorm variants [zhang2019root], and scaled dot-product attention [vaswani2017attention]. Efficient systems save or recompute only the tensors needed by these VJPs; checkpointing trades extra forward compute for lower activation memory [griewank2008]. |
C.3 Teach it
The one-sentence version. Backprop is shape-preserving bookkeeping: each operation receives an upstream gradient and returns one gradient per input.
An analogy. A VJP is an expense report. The loss sends one bill downstream; each operation splits that bill among the inputs that caused it.
At the board.
-
Write and derive the two matmul VJPs by summing over the repeated index.
-
Broadcast a bias across a batch, then sum the batch axis to get the bias gradient.
-
Show softmax as "subtract expected upstream under ".
-
Finish with attention as matmul, softmax, matmul in reverse.
Misconceptions to address. A gradient is not allowed to keep the output shape if the input shape was different. Broadcasting in the forward pass means summing in the backward pass. Softmax does not need a dense Jacobian.
Check for understanding. If a bias is added to every row of , why is a sum over rows?
C.4 Exercises
For with and , derive the VJP for .
Starting from , derive and .
Show that the softmax VJP can be computed as without forming the Jacobian.
Use the chapter code to compute the query VJP of scaled dot-product attention and finite-difference check it on a tiny seeded tensor.
References
-
[parr2018] T. Parr and J. Howard. The matrix calculus you need for deep learning. 2018. arXiv:1802.01528
-
[griewank2008] A. Griewank and A. Walther. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, 2nd edition. SIAM, 2008.
-
[ba2016layer] J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer Normalization. 2016. arXiv:1607.06450
-
[baydin2015automatic] A. G. Baydin et al. Automatic Differentiation in Machine Learning: a Survey. 2015. arXiv:1502.05767
-
[vaswani2017attention] A. Vaswani et al. Attention Is All You Need. 2017. arXiv:1706.03762
-
[zhang2019root] B. Zhang and R. Sennrich. Root Mean Square Layer Normalization. 2019. arXiv:1910.07467