Chapter 33

Supervised Fine-Tuning & LoRA

Chat templates, loss masking, packing, and low-rank adaptation.

Post-training begins by turning a pretrained next-token predictor into a model that answers in the format people expect. Supervised fine-tuning (SFT) does that with demonstrations: prompts plus ideal assistant replies. It is still ordinary next-token training, but now the data distribution is "a conversation followed by the answer we want." The details are small but unforgiving: role markers must match inference, the loss must ignore prompt tokens, packed sequences must not cross documents, and adapter weights must be cheap enough to train often.

33.1 Data, templates, and masks

An SFT example is a list of messages, not just a string. A chat template serializes it with role markers so the same tokens are seen during training and inference:

Listing 33.1 Chat templates and assistant-only targets
def render_chat(messages, add_generation_prompt=False):
    """Render role-marked messages into a tiny chat template string."""
    parts = []
    for message in messages:
        role = message["role"]
        if role not in ROLE_TOKENS:
            raise ValueError(f"unknown role: {role}")
        parts.extend([ROLE_TOKENS[role], message["content"], END_TOKEN])
    if add_generation_prompt:
        parts.append(ROLE_TOKENS["assistant"])
    return " ".join(parts)


def next_token_training_arrays(tokens, token_roles):
    """Inputs, next-token targets, and a mask for assistant targets only."""
    tokens = np.asarray(tokens)
    roles = np.asarray(token_roles)
    if tokens.shape != roles.shape:
        raise ValueError("tokens and token_roles must have the same shape")
    inputs = tokens[:-1]
    targets = tokens[1:]
    train_mask = roles[1:] == "assistant"
    return inputs, targets, train_mask

The template is part of the model contract. If demonstrations use <|assistant|> but production uses a different prefix, the first answer token is conditioned on a pattern the model did not practice. A real tokenizer would turn the rendered string into token ids; the code keeps the example text-level so the masking rule is visible.

If the target sequence is y1,…,yTy_1,\ldots,y_T and mt∈{0,1}m_t \in \{0,1\} marks assistant targets, the SFT loss is

L=−1M∑t=1Tmtlog⁡pθ(yt∣y<t),M=∑tmt.(33.1)L = -\frac{1}{M}\sum_{t=1}^{T} m_t \log p_\vtheta(y_t \mid y_{<t}), \qquad M = \sum_t m_t .\tag{33.1}

The mask is the difference between instruction tuning and training the model to imitate the user. Prompt tokens remain in the context, so they can influence the answer, but their own next-token predictions carry zero weight. For logits zt\vz_t and probabilities pt=softmax⁡(zt)\vp_t=\softmax(\vz_t), differentiating the log-softmax gives

zˉt,k=mtM(pt,k−1[k=yt]).(33.2)\bar{z}_{t,k} = \frac{m_t}{M}\big(p_{t,k} - \one[k=y_t]\big).\tag{33.2}

The test for this chapter gradient-checks that expression and also perturbs masked-out prompt logits to prove the scalar loss is unchanged.

The derivation is the same as pretraining, with one multiplier. Since ∂log⁡pt,y/∂zt,k=1[k=y]−pt,k\partial \log p_{t,y}/\partial z_{t,k}=\one[k=y]-p_{t,k}, the negative log-likelihood contributes pt,k−1[k=yt]p_{t,k}-\one[k=y_t]. Averaging only over assistant targets divides by MM and multiplying by mtm_t erases every prompt position. Whether to include the assistant end-of-turn token is a data-policy choice; the key is that the mask says exactly which targets represent desired assistant behavior.

Listing 33.2 Masked next-token cross-entropy
def softmax(logits, axis=-1):
    """Stable softmax."""
    shifted = logits - np.max(logits, axis=axis, keepdims=True)
    exp = np.exp(shifted)
    return exp / np.sum(exp, axis=axis, keepdims=True)


def masked_cross_entropy(logits, targets, train_mask):
    """Mean next-token cross-entropy over masked positions, with gradient."""
    logits = np.asarray(logits)
    targets = np.asarray(targets, dtype=np.int64)
    mask = np.asarray(train_mask, dtype=bool)
    if not np.any(mask):
        raise ValueError("at least one position must be trainable")
    probabilities = softmax(logits, axis=-1)
    rows = np.arange(targets.shape[0])
    losses = -np.log(probabilities[rows, targets])
    normalizer = np.sum(mask)
    loss = np.sum(np.where(mask, losses, 0.0)) / normalizer
    grad_logits = probabilities.copy()
    grad_logits[rows, targets] -= 1.0
    grad_logits *= mask[:, None] / normalizer
    return loss, grad_logits

33.2 Packing without accidental examples

Short conversations waste memory if each occupies a full context window. Sequence packing concatenates multiple documents into one fixed-length row, then runs the same causal model on the row. The trap is a fake training pair at each seam: the last token of one document should not predict the first token of the next.

The safe pattern is to insert a boundary token and carry a document-id array beside the token array. A next-token position is trainable only when source and target belong to the same real document.

This is separate from attention masking. A model may be allowed to attend across packed examples for efficiency, or it may receive a block-diagonal attention mask; either way, the loss must not reward cross-document predictions. The minimal invariant is that every target with a boundary or padding document id has zero loss weight.

Listing 33.3 Packing documents with boundary-aware masks
def packed_lm_arrays(packed_tokens, packed_doc_ids):
    """Next-token arrays whose mask forbids document-boundary targets."""
    inputs = packed_tokens[:, :-1]
    targets = packed_tokens[:, 1:]
    same_document = packed_doc_ids[:, :-1] == packed_doc_ids[:, 1:]
    train_mask = same_document & (packed_doc_ids[:, 1:] >= 0)
    return inputs, targets, train_mask

This keeps the compute benefit of packing while preserving the data distribution. Padding and boundary targets get document id −1-1, so their loss mask is false. A long document may still be split across rows; then the missing transition is dropped rather than replaced by a false cross-document transition.

Packing also makes validation stricter. If two documents are accidentally joined without a boundary, the model sees a fluent but bogus continuation and the loss says it is correct. The boundary id and the document-id mask give the tests something concrete to assert: no true mask entry is allowed where adjacent ids differ.

33.3 Low-rank adaptation

Full fine-tuning updates every matrix W∈Rdin×dout\mW \in \R^{d_{in}\times d_{out}}. LoRA keeps W\mW frozen and trains a low-rank update [hu2021lora]:

W′=W+αrBA,A∈Rr×dout,  B∈Rdin×r.(33.3)\mW' = \mW + \frac{\alpha}{r}\mB\mA, \qquad \mA \in \R^{r\times d_{out}},\; \mB \in \R^{d_{in}\times r}.\tag{33.3}

The code initializes B=0\mB=0, so W′=W\mW'=\mW and the adapted layer has identical outputs at step 0. A\mA is random; once B\mB moves, both factors can cooperate. After training, the update is merged into W\mW for inference, so serving has no extra matrix multiply.

The rank rr controls the adapter’s capacity. Rank 1 can only add an outer product; larger ranks add a sum of such products. The scale α/r\alpha/r keeps update magnitudes comparable when rr changes, so changing the rank does not automatically change the first useful step size. At initialization Aˉ=0\bar{\mA}=0 because B=0\mB=0, but Bˉ\bar{\mB} is generally nonzero, so the adapter starts moving immediately without changing the initial model output.

Listing 33.4 LoRA initialization, forward pass, and merge
def init_lora(d_in, d_out, rank, rng, scale=0.01):
    """Return A and zero B so the LoRA layer matches the frozen base at init."""
    a = rng.normal(0.0, scale, size=(rank, d_out)).astype(np.float32)
    b = np.zeros((d_in, rank), dtype=np.float32)
    return a, b


def lora_output(x, w, a, b, alpha):
    """Compute X @ (W + (alpha / rank) * B @ A)."""
    rank = a.shape[0]
    adapted = w + (alpha / rank) * (b @ a)
    return x @ adapted


def merge_lora(w, a, b, alpha):
    """Fold the trained low-rank update into W for inference."""
    return w + (alpha / a.shape[0]) * (b @ a)

Let Δˉ\bar{\boldsymbol{\Delta}} be the gradient with respect to the effective update Δ=BA\boldsymbol{\Delta}=\mB\mA. From dL=⟨Δˉ,dBA+BdA⟩dL = \langle \bar{\boldsymbol{\Delta}}, d\mB\mA + \mB d\mA \rangle,

Bˉ=αrΔˉA⊤,Aˉ=αrB⊤Δˉ.(33.4)\bar{\mB}=\frac{\alpha}{r}\bar{\boldsymbol{\Delta}}\mA^\T, \qquad \bar{\mA}=\frac{\alpha}{r}\mB^\T\bar{\boldsymbol{\Delta}}.\tag{33.4}
Listing 33.5 LoRA gradients and parameter counting
def lora_gradients(x, a, b, alpha, grad_y):
    """Backpropagate through the LoRA update for a loss on Y."""
    scale = alpha / a.shape[0]
    grad_update = x.T @ grad_y
    grad_a = scale * b.T @ grad_update
    grad_b = scale * grad_update @ a.T
    return grad_a, grad_b


def lora_mse_loss_and_grads(x, w, a, b, alpha, target):
    """Tiny objective used by the chapter tests."""
    y = lora_output(x, w, a, b, alpha)
    diff = y - target
    loss = 0.5 * np.mean(diff * diff)
    grad_y = diff / diff.size
    grad_a, grad_b = lora_gradients(x, a, b, alpha, grad_y)
    return loss, grad_a, grad_b


def lora_parameter_savings(d_in, d_out, rank):
    base = d_in * d_out
    trainable = rank * (d_in + d_out)
    return base, trainable, base / trainable

The savings are immediate. A 4096×40964096\times4096 matrix has 16,777,216 weights; rank-8 LoRA trains 8(4096+4096)=65,5368(4096+4096)=65{,}536 weights, a 256x reduction, computed and asserted by the tests. The base weights still dominate memory during training.

The gradient formulas are also a useful shape check. Bˉ\bar{\mB} has the same shape as B\mB because ΔˉA⊤\bar{\boldsymbol{\Delta}}\mA^\T is din×rd_{in}\times r. Aˉ\bar{\mA} has the same shape as A\mA because B⊤Δˉ\mB^\T\bar{\boldsymbol{\Delta}} is r×doutr\times d_{out}. A finite-difference check on both factors catches transposes, missing scales, and accidental updates to the frozen matrix.

33.4 QLoRA in one paragraph

QLoRA stores the frozen base model in 4-bit NormalFloat (NF4) quantized blocks, dequantizes them for computation, and trains LoRA adapters on top [dettmers2023qlora]. Because the quantized base is frozen, gradients and optimizer state are needed only for the adapter weights. The price is quantization error and extra dequantization machinery; the benefit is that larger base models fit in the same accelerator memory.

Conceptually, QLoRA changes storage, not the supervised objective. The same chat template, assistant-only mask, packing boundaries, and LoRA merge logic still determine what the model learns.

In practice

Instruction-tuned systems use fixed templates for system, user, assistant, and tool messages; changing the template after training changes the model’s input distribution. The assistant-only loss is standard for demonstration data because the prompt is conditioning context, not behavior to imitate. LoRA adapters are commonly trained per task or customer and then merged or selected at inference, while QLoRA is useful when memory, not arithmetic, is the bottleneck [hu2021lora] [dettmers2023qlora]. RLHF pipelines often start from an SFT model before preference optimization [ouyang2022training].

Key equations
L=−1M∑tmtlog⁡pθ(yt∣y<t),M=∑tmtL = -\frac{1}{M}\sum_t m_t \log p_\vtheta(y_t \mid y_{<t}), \qquad M=\sum_t m_t
zˉt,k=mtM(pt,k−1[k=yt])\bar{z}_{t,k}=\frac{m_t}{M}\big(p_{t,k}-\one[k=y_t]\big)
W′=W+αrBA,B0=0⇒W0′=W\mW' = \mW + \frac{\alpha}{r}\mB\mA, \qquad \mB_0=0 \Rightarrow \mW'_0=\mW
Bˉ=αrΔˉA⊤,Aˉ=αrB⊤Δˉ\bar{\mB}=\frac{\alpha}{r}\bar{\boldsymbol{\Delta}}\mA^\T, \qquad \bar{\mA}=\frac{\alpha}{r}\mB^\T\bar{\boldsymbol{\Delta}}
LoRA parameters=r(din+dout)≪dindout\text{LoRA parameters}=r(d_{in}+d_{out}) \ll d_{in}d_{out}

33.5 Teach it

The one-sentence version. SFT teaches the model which replies to write, and LoRA makes that update low-rank so only a small adapter is trained.

An analogy. The prompt is the exam question and the assistant message is the answer key. You let the student read the whole question, but you grade only the answer.

At the board.

  1. Draw a chat as role-marked tokens: system, user, assistant.

  2. Put mask zeros over prompt tokens and ones over assistant tokens; derive the softmax gradient with the mask multiplier.

  3. Pack two short documents, then cross out the seam so no target crosses it.

  4. Factor a large update matrix as BA\mB\mA; set B=0\mB=0 to start from the base model exactly.

Misconceptions to address.

  • "The model should learn to predict user prompts." Not in SFT; prompts condition the answer.

  • "Packing is just concatenation." It also needs boundary-aware loss masks.

  • "LoRA changes inference forever." It can be merged into the base matrix.

Check for understanding. Why does B=0\mB=0 make a LoRA layer identical to the frozen layer at initialization even when A\mA is random?

33.6 Exercises

Exercise 33.1 ★ Template discipline

Explain why training and inference must use the same role markers. Then say which tokens are context and which tokens are targets in a user-assistant SFT example.

Exercise 33.2 ★★ Masked gradient

Derive (33.2) from (33.1) and the softmax derivative. What happens to the gradient at positions with mt=0m_t=0?

Exercise 33.3 ★★ Packed boundaries

Given two tokenized documents [1,2][1,2] and [3][3] with boundary token 99 and padding token 0, pack them into length-4 rows. Write the next-token targets and the Boolean loss mask.

Exercise 33.4 ★★★ LoRA implementation

For Y=X(W+αrBA)Y=X(\mW+\frac{\alpha}{r}\mB\mA), derive the gradients for A\mA and B\mB. Then use the chapter code to compute the parameter saving for din=dout=4096d_{in}=d_{out}=4096 and r=8r=8.

References

  • [dettmers2023qlora] T. Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs. 2023. arXiv:2305.14314

  • [hu2021lora] E. J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. 2021. arXiv:2106.09685

  • [ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155