Chapter 22

The Transformer Block

Pre-norm residual blocks, RMSNorm, SwiGLU feed-forward layers, and the decoder-only stack.

A decoder-only transformer is a residual stream repeatedly edited by attention and a feed-forward network. The block is small enough to write in NumPy, but expressive enough that stacking it gives the backbone of a GPT. This chapter uses the modern pre-norm form: normalize, transform, add back. It is still the same idea introduced by the Transformer architecture [vaswani2017attention], but with the decoder-only choices that make next-token prediction simple.

22.1 The pre-norm decoder block

Let xℓ\vx_\ell be the residual stream entering layer ℓ\ell. A pre-norm decoder block updates it in two residual steps:

uℓ=xℓ+Attn⁡(RMSNorm⁡(xℓ)),xℓ+1=uℓ+FFN⁡(RMSNorm⁡(uℓ)).(22.1)\begin{aligned} \vu_\ell &= \vx_\ell + \operatorname{Attn}(\operatorname{RMSNorm}(\vx_\ell)), \\ \vx_{\ell+1} &= \vu_\ell + \operatorname{FFN}(\operatorname{RMSNorm}(\vu_\ell)). \end{aligned}\tag{22.1}

The causal attention is multi-head attention with a triangular mask and RoPE from Chapter 21. The residual additions matter as much as the sublayers: they give each layer permission to make an incremental edit instead of rewriting the representation. In the residual-stream view, the vector at each token is a shared workspace; attention copies information between positions, and the feed-forward network transforms each position independently. This view is useful when debugging models. If a token needs a fact from an earlier token, the attention update can move that fact into the current token’s stream. If the token already has enough information, the residual path can carry it forward unchanged. The feed-forward update then acts like a learned table of local features: it can turn combinations of stream coordinates into a new direction that later heads or the output classifier will read. The backward pass follows the same map in reverse. The gradient first splits across the final residual add, then flows through the feed-forward branch and the identity branch. It splits again at the attention residual add. That repeated identity path is why very deep pre-norm stacks are easier to optimize than post-norm stacks that put a normalization after the addition.

Listing 22.1 The block in NumPy
def transformer_block_forward(x, params, n_heads, rope_base=10_000.0):
    attn_in, norm1 = rmsnorm_forward(x, params["attn_norm"])
    attn_out, attn_cache = causal_self_attention_forward(
        attn_in, params, n_heads, rope_base
    )
    residual = x + attn_out
    ffn_in, norm2 = rmsnorm_forward(residual, params["ffn_norm"])
    ffn_out, ffn_cache = swiglu_forward(ffn_in, params)
    out = residual + ffn_out
    return out, (norm1, attn_cache, residual, norm2, ffn_cache)

22.2 RMSNorm and SwiGLU

RMSNorm scales a vector by its root mean square and a learned gain [zhang2019root]:

RMSNorm⁡(x)i=gixi(1d∑j=1dxj2+ϵ)−1/2.(22.2)\operatorname{RMSNorm}(\vx)_i = g_i x_i \left(\frac{1}{d}\sum_{j=1}^{d}x_j^2+\epsilon\right)^{-1/2} .\tag{22.2}

Unlike LayerNorm, it does not subtract the mean. Its backward pass is one rank-one correction: if zˉ\bar{\vz} is the gradient after multiplying by the gain, rr is the inverse RMS, and c=∑izˉixic=\sum_i \bar{z}_i x_i, then

xˉi=rzˉi−xir3c/d.(22.3)\bar{x}_i = r\bar{z}_i - x_i r^3 c / d .\tag{22.3}

The feed-forward network here is the SwiGLU variant [shazeer2020glu]:

FFN⁡(x)=(SiLU⁡(xWg)⊙xWu)Wd.(22.4)\operatorname{FFN}(\vx)= (\operatorname{SiLU}(\vx\mW_g) \odot \vx\mW_u)\mW_d .\tag{22.4}

One projection makes gates, one makes values, and the down projection returns to the model width. This is a position-wise MLP; all tokens share the same weights. The gate is important because it lets the network choose which coordinates of the up projection are active for this token. SiLU is smooth, so a nearly closed gate still has a gradient. In the manual backward pass, the up path receives the gate value, while the gate path receives the up value times the derivative of SiLU. That symmetry is what makes the implementation short enough to gradient-check directly.

Listing 22.2 RMSNorm forward and backward
def rmsnorm_forward(x, weight, eps=1e-5):
    mean_square = np.mean(x * x, axis=-1, keepdims=True)
    inv_rms = 1.0 / np.sqrt(mean_square + eps)
    normalized = x * inv_rms
    return normalized * weight, (x, weight, inv_rms, normalized)


def rmsnorm_backward(dout, cache):
    x, weight, inv_rms, normalized = cache
    dnormalized = dout * weight
    scale_grad = np.sum(dnormalized * x, axis=-1, keepdims=True)
    width = x.shape[-1]
    dx = inv_rms * dnormalized - x * inv_rms ** 3 * scale_grad / width
    dweight = np.sum(dout * normalized, axis=tuple(range(dout.ndim - 1)))
    return dx, dweight
Listing 22.3 SwiGLU feed-forward
def swiglu_forward(x, params):
    gate = x @ params["W_gate"]
    up = x @ params["W_up"]
    hidden = silu(gate) * up
    out = hidden @ params["W_down"]
    return out, (x, params, gate, up, hidden)

22.3 Stacking, logits, and weight tying

A tiny GPT begins with token IDs, looks up rows of an embedding matrix E∈RV×d\mE \in \R^{V\times d}, passes the sequence through LL blocks, applies a final normalization, and produces logits. With weight tying [press2016using], the output classifier reuses the embedding matrix:

logits⁡t=htE⊤.(22.5)\operatorname{logits}_t = \vh_t \mE^{\T} .\tag{22.5}

The tied head saves parameters and keeps input and output token spaces aligned. Training applies a softmax cross-entropy at every position against the next token. At inference time, the same stack runs on the prefix and samples the next token from the last-position logits. The embedding lookup has a sparse backward pass: only rows whose token IDs appeared in the batch receive input-side gradients. With tying, the same matrix also receives dense output-side gradients from the classifier. The next chapter uses both contributions in one NumPy training loop; no special framework feature is required, only careful accumulation into shared rows.

22.4 Parameters and FLOPs

Ignore biases and normalization gains first. Attention has four dense d×dd\times d matrices: WQ,WK,WV,WO\mW_Q,\mW_K,\mW_V,\mW_O, so it has 4d24d^2 parameters. A plain two-layer MLP with hidden width 4d4d has d(4d)+(4d)d=8d2d(4d)+(4d)d=8d^2 parameters, giving the familiar 12d212d^2 per block. The SwiGLU code uses hidden width hh, so its feed-forward count is 3dh3dh and the block has 4d2+3dh4d^2+3dh parameters; choosing h=4dh=4d makes it 16d216d^2, while h=8d/3h=8d/3 keeps the block near 12d212d^2.

Counting a multiply-add as two FLOPs, dense forward compute is about 2P2P FLOPs per token for PP active parameters. Backpropagation computes gradients with respect to activations and weights, so training is roughly three times the forward dense cost, or 6P6P FLOPs per token. Over DD training tokens and NN model parameters, this gives the useful preview C≈6NDC \approx 6ND [hoffmann2022training]. Full attention also adds about 4Td4Td FLOPs per token per layer for the score and value mixing over a context of length TT. For short contexts and wide models, the dense matrices dominate. For very long contexts, the TT term becomes visible, which is why later chapters care about KV caches and efficient attention kernels. The rule 6ND6ND is therefore a planning approximation, not a profiler: it ignores embeddings, normalization, optimizer overhead, and hardware utilization. Still, it is a powerful mental model: doubling parameters or doubling tokens roughly doubles the dense training compute, so architecture choices that change PP matter immediately in budget planning for every serious training run.

In practice

Modern decoder blocks are usually pre-norm, because gradients can flow along the residual stream before entering a sublayer; analyses of normalization placement explain why this stabilizes deep Transformers [xiong2020layer]. RMSNorm, SwiGLU, and RoPE appear together in influential open-weight decoder families such as LLaMA [touvron2023llama]. The exact feed-forward width is a budget choice: SwiGLU with h=4dh=4d is larger than a classic 4d4d MLP, so many model families reduce hh when matching a parameter budget.

Key equations
uℓ=xℓ+Attn⁡(RMSNorm⁡(xℓ))\vu_\ell = \vx_\ell + \operatorname{Attn}(\operatorname{RMSNorm}(\vx_\ell))
xℓ+1=uℓ+FFN⁡(RMSNorm⁡(uℓ))\vx_{\ell+1}=\vu_\ell+\operatorname{FFN}(\operatorname{RMSNorm}(\vu_\ell))
FFN⁡(x)=(SiLU⁡(xWg)⊙xWu)Wd\operatorname{FFN}(\vx)=(\operatorname{SiLU}(\vx\mW_g)\odot\vx\mW_u)\mW_d
Pblock≈4d2+3dhP_{\text{block}} \approx 4d^2 + 3dh
Ctrain≈6NDC_{\text{train}} \approx 6ND

22.5 Teach it

The one-sentence version. A transformer block normalizes the residual stream, lets attention move information across positions, adds the result back, then normalizes again and applies a gated MLP at each position.

An analogy. The residual stream is a shared document. Attention copies notes between paragraphs; the feed-forward network rewrites each paragraph locally; residual connections keep the original text unless an edit is useful.

At the board.

  1. Draw the two residual arrows in (22.1).

  2. Write RMSNorm as "scale by inverse RMS, then by a learned gain."

  3. Expand SwiGLU into gate, up, elementwise product, down.

  4. Count matrices: four attention matrices and three SwiGLU matrices.

Misconceptions to address.

  • "The block output is only the attention output." The residual stream always carries through.

  • "SwiGLU is just a bigger ReLU MLP." It gates one projection by a smooth function of another.

  • "The output head must be separate." With weight tying, it is the embedding matrix transposed.

Check for understanding. If h=4dh=4d, why does a SwiGLU feed-forward have more parameters than the classic 4d4d two-layer MLP?

22.6 Exercises

Exercise 22.1 ★ Residual-stream view

Explain in words what attention and the feed-forward network each contribute to the residual stream in a decoder block. Why does pre-norm help the residual path stay direct?

Exercise 22.2 ★★ RMSNorm backward

Derive (22.3) from zi=xirz_i=x_i r with r=(d−1∑jxj2+ϵ)−1/2r=(d^{-1}\sum_j x_j^2+\epsilon)^{-1/2}, ignoring the learned gain until the last step.

Exercise 22.3 ★★ Parameter count

For model width dd and SwiGLU hidden width hh, count the attention and feed-forward parameters in one block. Evaluate the formula for h=4dh=4d and for h=8d/3h=8d/3.

Exercise 22.4 ★★★ Gradient-check a block

Using a one-layer, tiny-width block, define a scalar loss ∑y⊙yˉ\sum y \odot \bar{y} for a fixed upstream array yˉ\bar{y}. Check the gradients with respect to both the input and all block parameters by finite differences.

References

  • [hoffmann2022training] J. Hoffmann et al. Training Compute-Optimal Large Language Models. 2022. arXiv:2203.15556

  • [press2016using] O. Press and L. Wolf. Using the Output Embedding to Improve Language Models. 2016. arXiv:1608.05859

  • [shazeer2020glu] N. Shazeer. GLU Variants Improve Transformer. 2020. arXiv:2002.05202

  • [touvron2023llama] H. Touvron et al. LLaMA: Open and Efficient Foundation Language Models. 2023. arXiv:2302.13971

  • [vaswani2017attention] A. Vaswani et al. Attention Is All You Need. 2017. arXiv:1706.03762

  • [xiong2020layer] R. Xiong et al. On Layer Normalization in the Transformer Architecture. 2020. arXiv:2002.04745

  • [zhang2019root] B. Zhang and R. Sennrich. Root Mean Square Layer Normalization. 2019. arXiv:1910.07467