Chapter 21

Positional Encoding & RoPE

Sinusoidal and learned positions, rotary embeddings, ALiBi, and context extension with YaRN.

Self-attention sees a set of token vectors unless we tell it where each token came from. That is fatal for language: dog bites man and man bites dog contain the same embeddings but mean different things. Positional encodings inject order while keeping the attention calculation small, and RoPE is the relative-position form used by many current decoder-only LLMs.

21.1 Attention without positions

Scaled dot-product attention is permutation-equivariant. Let P\mP be a permutation matrix that reorders the sequence rows. With Q=XWQ\mQ = \mX\mW_Q, and similarly for keys and values,

Attn⁡(PQ,PK,PV)=PAttn⁡(Q,K,V).(21.1)\operatorname{Attn}(\mP\mQ,\mP\mK,\mP\mV) = \mP\operatorname{Attn}(\mQ,\mK,\mV) .\tag{21.1}

The reason is mechanical: scores become PQK⊤P⊤\mP\mQ\mK^{\T}\mP^{\T}, rowwise softmax only permutes rows and columns, and multiplying by PV\mP\mV permutes the outputs. This is useful for sets, but a decoder needs order. The test below permutes inputs and verifies that the outputs permute in exactly the same way. The point is not that the layer is weak; it can still compare every token with every other token. The missing fact is which occurrence came first. If we give attention only the content vector for the, every identical the has the same raw representation before context is mixed. A positional signal breaks that symmetry while leaving the rest of the attention machinery unchanged.

Listing 21.1 Permutation equivariance of attention with no positions
def permutation_error(x, wq, wk, wv, permutation):
    q, k, v = x @ wq, x @ wk, x @ wv
    original = attention(q, k, v)
    xp = x[permutation]
    permuted = attention(xp @ wq, xp @ wk, xp @ wv)
    return np.max(np.abs(permuted - original[permutation]))

21.2 Absolute positions

The simplest fix is to add a position vector before forming queries, keys, and values: xt←xt+pt\vx_t \leftarrow \vx_t + \vp_t. In the original Transformer, pt\vp_t was a fixed sinusoidal table [vaswani2017attention]:

pt,2i=sin⁡(tθi),pt,2i+1=cos⁡(tθi),θi=b−2i/d.(21.2)p_{t,2i}=\sin(t\theta_i), \quad p_{t,2i+1}=\cos(t\theta_i), \quad \theta_i=b^{-2i/d} .\tag{21.2}

The sine and cosine make nearby positions smooth and give each two-dimensional pair a different wavelength. A learned absolute table uses the same addition but stores pt\vp_t as parameters. It is flexible inside the trained context length, but extrapolating beyond the last learned row has no natural meaning. Absolute encodings are easy to reason about: every layer receives content plus a coordinate. Their cost is that attention scores see positions only through the learned projections applied after addition. The model must learn from data how two absolute coordinates imply a relative distance such as "the previous token" or "the matching opener many steps ago." Sinusoidal tables help because their phases shift predictably, while learned tables spend parameters for maximum freedom over the training window. The addition is also global: once pt\vp_t is added, every downstream linear map mixes content and coordinate together. That is fine for short trained lengths, but it makes it hard to separate "what token is this?" from "where did it appear?" when asking why a long-context extrapolation failed.

Listing 21.2 Absolute position tables
def sinusoidal_positions(length, dim, base=10_000.0):
    """The original fixed absolute table: sin on even dims, cos on odd dims."""
    if dim % 2:
        raise ValueError("sinusoidal positions need an even dimension")
    positions = np.arange(length, dtype=np.float64)[:, None]
    theta = base ** (-np.arange(0, dim, 2, dtype=np.float64) / dim)
    angles = positions * theta[None, :]
    table = np.empty((length, dim), dtype=np.float64)
    table[:, 0::2] = np.sin(angles)
    table[:, 1::2] = np.cos(angles)
    return table


def learned_positions(length, dim, rng, scale=0.02):
    """A learned absolute position table, initialized like a small embedding."""
    return rng.normal(0.0, scale, size=(length, dim)).astype(np.float32)

21.3 Rotary positions

RoPE puts position into the queries and keys, not the values. Split a head into adjacent pairs. For pair ii, rotate both the query and key at position tt by angle tθit\theta_i, where θi=b−2i/d\theta_i=b^{-2i/d}. The two-dimensional rotation is

Rt(i)=[cos⁡(tθi)−sin⁡(tθi)sin⁡(tθi)cos⁡(tθi)].(21.3)R_t^{(i)} = \begin{bmatrix} \cos(t\theta_i) & -\sin(t\theta_i) \\ \sin(t\theta_i) & \cos(t\theta_i) \end{bmatrix} .\tag{21.3}

The payoff is relative. Since rotations are orthogonal and angles add, Rm⊤Rn=Rn−mR_m^{\T}R_n = R_{n-m}, so

⟨Rmq,Rnk⟩=q⊤Rn−mk.(21.4)\langle R_m\vq, R_n\vk\rangle = \vq^{\T} R_{n-m}\vk .\tag{21.4}

The attention score therefore depends on the token contents and the offset n−mn-m, not on the absolute positions separately [su2021roformer]. The implementation is just elementwise cosines and sines; the backward pass through RoPE is the inverse rotation. RoPE also preserves the norm of every query and key pair, because a rotation is orthogonal. That means it changes the angle used in the dot product without changing the scale that the softmax sees. Different pairs rotate at different frequencies, so short and long offsets leave different fingerprints across the head dimension. Values are usually left unrotated: once attention decides which positions to read from, the value stream should carry content rather than another copy of the coordinate system. This separation is why RoPE fits naturally inside multi-head attention: each head can learn content projections, then receive the same deterministic relative geometry. The learned weights decide how much to use that geometry.

Listing 21.3 Rotating adjacent pairs with cosines and sines
def apply_rope(x, positions, base=10_000.0, inverse=False):
    """Rotate each adjacent 2-D pair by positions * theta_i."""
    x = np.asarray(x)
    if x.shape[-1] % 2:
        raise ValueError("RoPE needs an even last dimension")
    positions = np.asarray(positions, dtype=x.dtype)
    if positions.ndim == 1 and x.ndim > 2:
        shape = [1] * (x.ndim - 1)
        shape[1] = positions.size
        positions = positions.reshape(shape)
    theta = rope_frequencies(x.shape[-1], base).astype(x.dtype)
    angles = positions[..., None] * theta
    if inverse:
        angles = -angles
    cos, sin = np.cos(angles), np.sin(angles)
    y = np.empty_like(x)
    even, odd = x[..., 0::2], x[..., 1::2]
    y[..., 0::2] = even * cos - odd * sin
    y[..., 1::2] = even * sin + odd * cos
    return y

21.4 Biases and longer contexts

ALiBi does not rotate or add vectors. It adds a head-specific linear penalty to each attention score, so a query at tt attending to an earlier key ss receives −αh(t−s)-\alpha_h(t-s) [press2021train].

Listing 21.4 A linear distance bias for causal attention
def alibi_bias(length, slope):
    """Causal ALiBi scores: query t gets -slope * (t - s) for key s <= t."""
    t = np.arange(length)[:, None]
    s = np.arange(length)[None, :]
    return -float(slope) * np.maximum(t - s, 0)

NoPE is the ablation that drops explicit position features and lets the causal mask and data statistics carry order, which is rarely the safest default. For extending RoPE contexts, position interpolation evaluates a long context through compressed positions t′=t/Lt' = t/L [chen2023extending]. NTK-aware scaling changes the RoPE base so low-frequency pairs stretch farther while high-frequency pairs keep local resolution [bloc972023]. YaRN combines scaled frequencies with a short fine-tuning recipe for efficient context extension [peng2023yarn]. All three extension tricks are compromises. Compressing positions makes the model reuse the angles it saw during training, but it also makes nearby long-context tokens look closer together than before. Changing the base preserves the formula while moving the frequency grid. YaRN adds a practical recipe around that idea rather than claiming the original model learned the new length from scratch.

In practice

Decoder-only LLMs commonly prefer relative or rotary schemes because the score for a pair of tokens can depend directly on their distance, not just on two absolute IDs. RoPE is standard in many open-weight transformer families, while ALiBi is attractive when length extrapolation is the main goal [press2021train]. Context extension methods should be treated as compatibility patches: they can make a model run at longer lengths, but they do not create training data at those lengths. When changing context length, test the task distribution itself: retrieval, summarization, and code completion can stress different offsets even when they use the same maximum sequence length.

Key equations
Attn⁡(Q,K,V)=softmax⁡(QK⊤/dh)V\operatorname{Attn}(\mQ,\mK,\mV) = \softmax(\mQ\mK^{\T}/\sqrt{d_h})\mV
pt,2i=sin⁡(tθi),pt,2i+1=cos⁡(tθi)p_{t,2i}=\sin(t\theta_i), \quad p_{t,2i+1}=\cos(t\theta_i)
θi=b−2i/d\theta_i=b^{-2i/d}
⟨Rmq,Rnk⟩=q⊤Rn−mk\langle R_m\vq, R_n\vk\rangle = \vq^{\T}R_{n-m}\vk
ALiBi⁡h,t,s=−αh(t−s)(s≤t)\operatorname{ALiBi}_{h,t,s}=-\alpha_h(t-s) \quad (s\le t)

21.5 Teach it

The one-sentence version. Attention needs positions because otherwise reordering the tokens only reorders the outputs; RoPE gives each query-key dot product a relative offset.

An analogy. Absolute positions are seat numbers on tickets. RoPE is more like turning your head by an amount that depends on where you sit, so two people compare directions by the angle between them.

At the board.

  1. Write softmax⁡(QK⊤)V\softmax(\mQ\mK^{\T})\mV, then wrap every matrix in the same permutation and cancel it to show equivariance.

  2. Draw one adjacent pair (x0,x1)(x_0,x_1) and rotate it by tθit\theta_i.

  3. Multiply Rm⊤RnR_m^{\T}R_n and point to the angle (n−m)θi(n-m)\theta_i.

  4. Add the ALiBi line: a fixed negative slope as keys get farther into the past.

Misconceptions to address.

  • "The position vector stores the word order by itself." It only gives the network a coordinate.

  • "RoPE changes values." Standard RoPE rotates queries and keys; values stay in content space.

  • "Long-context scaling is free." It changes the geometry a model was trained on.

Check for understanding. If every token vector and every position vector were permuted together, would absolute-position attention notice the original order?

21.6 Exercises

Exercise 21.1 ★ Prove equivariance

Prove (21.1) by following the score matrix, the rowwise softmax, and the final value multiplication. Why is this bad for a language model with no positional signal?

Exercise 21.2 ★★ RoPE is relative

Starting from the two-dimensional rotation matrix, prove (21.4). Then explain why shifting both positions by the same amount leaves the RoPE dot product unchanged.

Exercise 21.3 ★★ ALiBi row

For a causal sequence of length LL, write the ALiBi bias row for the last query in terms of the slope α\alpha. What happens when α\alpha is larger?

Exercise 21.4 ★★★ Implement and test RoPE

Write a vectorized RoPE function for inputs shaped (B,T,H,dh)(B,T,H,d_h). Check that applying the inverse rotation recovers the input, and that the dot product is unchanged when both positions are shifted equally.

References

  • [bloc972023] bloc97. NTK-aware scaled RoPE allows LLaMA models to have extended context size without fine-tuning. Reddit r/LocalLLaMA post, June 2023.

  • [chen2023extending] S. Chen et al. Extending Context Window of Large Language Models via Positional Interpolation. 2023. arXiv:2306.15595

  • [peng2023yarn] B. Peng et al. YaRN: Efficient Context Window Extension of Large Language Models. 2023. arXiv:2309.00071

  • [press2021train] O. Press, N. A. Smith, and M. Lewis. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. 2021. arXiv:2108.12409

  • [su2021roformer] J. Su et al. RoFormer: Enhanced Transformer with Rotary Position Embedding. 2021. arXiv:2104.09864

  • [vaswani2017attention] A. Vaswani et al. Attention Is All You Need. 2017. arXiv:1706.03762