Chapter 22
The Transformer Block
Pre-norm residual blocks, RMSNorm, SwiGLU feed-forward layers, and the decoder-only stack.
A decoder-only transformer is a residual stream repeatedly edited by attention and a feed-forward network. The block is small enough to write in NumPy, but expressive enough that stacking it gives the backbone of a GPT. This chapter uses the modern pre-norm form: normalize, transform, add back. It is still the same idea introduced by the Transformer architecture [vaswani2017attention], but with the decoder-only choices that make next-token prediction simple.
22.1 The pre-norm decoder block
Let be the residual stream entering layer . A pre-norm decoder block updates it in two residual steps:
The causal attention is multi-head attention with a triangular mask and RoPE from Chapter 21. The residual additions matter as much as the sublayers: they give each layer permission to make an incremental edit instead of rewriting the representation. In the residual-stream view, the vector at each token is a shared workspace; attention copies information between positions, and the feed-forward network transforms each position independently. This view is useful when debugging models. If a token needs a fact from an earlier token, the attention update can move that fact into the current token’s stream. If the token already has enough information, the residual path can carry it forward unchanged. The feed-forward update then acts like a learned table of local features: it can turn combinations of stream coordinates into a new direction that later heads or the output classifier will read. The backward pass follows the same map in reverse. The gradient first splits across the final residual add, then flows through the feed-forward branch and the identity branch. It splits again at the attention residual add. That repeated identity path is why very deep pre-norm stacks are easier to optimize than post-norm stacks that put a normalization after the addition.
def transformer_block_forward(x, params, n_heads, rope_base=10_000.0):
attn_in, norm1 = rmsnorm_forward(x, params["attn_norm"])
attn_out, attn_cache = causal_self_attention_forward(
attn_in, params, n_heads, rope_base
)
residual = x + attn_out
ffn_in, norm2 = rmsnorm_forward(residual, params["ffn_norm"])
ffn_out, ffn_cache = swiglu_forward(ffn_in, params)
out = residual + ffn_out
return out, (norm1, attn_cache, residual, norm2, ffn_cache)
22.2 RMSNorm and SwiGLU
RMSNorm scales a vector by its root mean square and a learned gain [zhang2019root]:
Unlike LayerNorm, it does not subtract the mean. Its backward pass is one rank-one correction: if is the gradient after multiplying by the gain, is the inverse RMS, and , then
The feed-forward network here is the SwiGLU variant [shazeer2020glu]:
One projection makes gates, one makes values, and the down projection returns to the model width.
This is a position-wise MLP; all tokens share the same weights.
The gate is important because it lets the network choose which coordinates of the up projection
are active for this token. SiLU is smooth, so a nearly closed gate still has a gradient. In the
manual backward pass, the up path receives the gate value, while the gate path receives the up
value times the derivative of SiLU. That symmetry is what makes the implementation short enough
to gradient-check directly.
def rmsnorm_forward(x, weight, eps=1e-5):
mean_square = np.mean(x * x, axis=-1, keepdims=True)
inv_rms = 1.0 / np.sqrt(mean_square + eps)
normalized = x * inv_rms
return normalized * weight, (x, weight, inv_rms, normalized)
def rmsnorm_backward(dout, cache):
x, weight, inv_rms, normalized = cache
dnormalized = dout * weight
scale_grad = np.sum(dnormalized * x, axis=-1, keepdims=True)
width = x.shape[-1]
dx = inv_rms * dnormalized - x * inv_rms ** 3 * scale_grad / width
dweight = np.sum(dout * normalized, axis=tuple(range(dout.ndim - 1)))
return dx, dweight
def swiglu_forward(x, params):
gate = x @ params["W_gate"]
up = x @ params["W_up"]
hidden = silu(gate) * up
out = hidden @ params["W_down"]
return out, (x, params, gate, up, hidden)
22.3 Stacking, logits, and weight tying
A tiny GPT begins with token IDs, looks up rows of an embedding matrix , passes the sequence through blocks, applies a final normalization, and produces logits. With weight tying [press2016using], the output classifier reuses the embedding matrix:
The tied head saves parameters and keeps input and output token spaces aligned. Training applies a softmax cross-entropy at every position against the next token. At inference time, the same stack runs on the prefix and samples the next token from the last-position logits. The embedding lookup has a sparse backward pass: only rows whose token IDs appeared in the batch receive input-side gradients. With tying, the same matrix also receives dense output-side gradients from the classifier. The next chapter uses both contributions in one NumPy training loop; no special framework feature is required, only careful accumulation into shared rows.
22.4 Parameters and FLOPs
Ignore biases and normalization gains first. Attention has four dense matrices: , so it has parameters. A plain two-layer MLP with hidden width has parameters, giving the familiar per block. The SwiGLU code uses hidden width , so its feed-forward count is and the block has parameters; choosing makes it , while keeps the block near .
Counting a multiply-add as two FLOPs, dense forward compute is about FLOPs per token for active parameters. Backpropagation computes gradients with respect to activations and weights, so training is roughly three times the forward dense cost, or FLOPs per token. Over training tokens and model parameters, this gives the useful preview [hoffmann2022training]. Full attention also adds about FLOPs per token per layer for the score and value mixing over a context of length . For short contexts and wide models, the dense matrices dominate. For very long contexts, the term becomes visible, which is why later chapters care about KV caches and efficient attention kernels. The rule is therefore a planning approximation, not a profiler: it ignores embeddings, normalization, optimizer overhead, and hardware utilization. Still, it is a powerful mental model: doubling parameters or doubling tokens roughly doubles the dense training compute, so architecture choices that change matter immediately in budget planning for every serious training run.
|
In practice
|
Modern decoder blocks are usually pre-norm, because gradients can flow along the residual stream before entering a sublayer; analyses of normalization placement explain why this stabilizes deep Transformers [xiong2020layer]. RMSNorm, SwiGLU, and RoPE appear together in influential open-weight decoder families such as LLaMA [touvron2023llama]. The exact feed-forward width is a budget choice: SwiGLU with is larger than a classic MLP, so many model families reduce when matching a parameter budget. |
22.5 Teach it
The one-sentence version. A transformer block normalizes the residual stream, lets attention move information across positions, adds the result back, then normalizes again and applies a gated MLP at each position.
An analogy. The residual stream is a shared document. Attention copies notes between paragraphs; the feed-forward network rewrites each paragraph locally; residual connections keep the original text unless an edit is useful.
At the board.
-
Draw the two residual arrows in (22.1).
-
Write RMSNorm as "scale by inverse RMS, then by a learned gain."
-
Expand SwiGLU into gate, up, elementwise product, down.
-
Count matrices: four attention matrices and three SwiGLU matrices.
Misconceptions to address.
-
"The block output is only the attention output." The residual stream always carries through.
-
"SwiGLU is just a bigger ReLU MLP." It gates one projection by a smooth function of another.
-
"The output head must be separate." With weight tying, it is the embedding matrix transposed.
Check for understanding. If , why does a SwiGLU feed-forward have more parameters than the classic two-layer MLP?
22.6 Exercises
Explain in words what attention and the feed-forward network each contribute to the residual stream in a decoder block. Why does pre-norm help the residual path stay direct?
Derive (22.3) from with , ignoring the learned gain until the last step.
For model width and SwiGLU hidden width , count the attention and feed-forward parameters in one block. Evaluate the formula for and for .
Using a one-layer, tiny-width block, define a scalar loss for a fixed upstream array . Check the gradients with respect to both the input and all block parameters by finite differences.
References
-
[hoffmann2022training] J. Hoffmann et al. Training Compute-Optimal Large Language Models. 2022. arXiv:2203.15556
-
[press2016using] O. Press and L. Wolf. Using the Output Embedding to Improve Language Models. 2016. arXiv:1608.05859
-
[shazeer2020glu] N. Shazeer. GLU Variants Improve Transformer. 2020. arXiv:2002.05202
-
[touvron2023llama] H. Touvron et al. LLaMA: Open and Efficient Foundation Language Models. 2023. arXiv:2302.13971
-
[vaswani2017attention] A. Vaswani et al. Attention Is All You Need. 2017. arXiv:1706.03762
-
[xiong2020layer] R. Xiong et al. On Layer Normalization in the Transformer Architecture. 2020. arXiv:2002.04745
-
[zhang2019root] B. Zhang and R. Sennrich. Root Mean Square Layer Normalization. 2019. arXiv:1910.07467