Chapter 33
Supervised Fine-Tuning & LoRA
Chat templates, loss masking, packing, and low-rank adaptation.
Post-training begins by turning a pretrained next-token predictor into a model that answers in the format people expect. Supervised fine-tuning (SFT) does that with demonstrations: prompts plus ideal assistant replies. It is still ordinary next-token training, but now the data distribution is "a conversation followed by the answer we want." The details are small but unforgiving: role markers must match inference, the loss must ignore prompt tokens, packed sequences must not cross documents, and adapter weights must be cheap enough to train often.
33.1 Data, templates, and masks
An SFT example is a list of messages, not just a string. A chat template serializes it with role markers so the same tokens are seen during training and inference:
def render_chat(messages, add_generation_prompt=False):
"""Render role-marked messages into a tiny chat template string."""
parts = []
for message in messages:
role = message["role"]
if role not in ROLE_TOKENS:
raise ValueError(f"unknown role: {role}")
parts.extend([ROLE_TOKENS[role], message["content"], END_TOKEN])
if add_generation_prompt:
parts.append(ROLE_TOKENS["assistant"])
return " ".join(parts)
def next_token_training_arrays(tokens, token_roles):
"""Inputs, next-token targets, and a mask for assistant targets only."""
tokens = np.asarray(tokens)
roles = np.asarray(token_roles)
if tokens.shape != roles.shape:
raise ValueError("tokens and token_roles must have the same shape")
inputs = tokens[:-1]
targets = tokens[1:]
train_mask = roles[1:] == "assistant"
return inputs, targets, train_mask
The template is part of the model contract. If demonstrations use <|assistant|> but production uses a different prefix, the first answer token is conditioned on a pattern the model did not practice. A real tokenizer would turn the rendered string into token ids; the code keeps the example text-level so the masking rule is visible.
If the target sequence is and marks assistant targets, the SFT loss is
The mask is the difference between instruction tuning and training the model to imitate the user. Prompt tokens remain in the context, so they can influence the answer, but their own next-token predictions carry zero weight. For logits and probabilities , differentiating the log-softmax gives
The test for this chapter gradient-checks that expression and also perturbs masked-out prompt logits to prove the scalar loss is unchanged.
The derivation is the same as pretraining, with one multiplier. Since , the negative log-likelihood contributes . Averaging only over assistant targets divides by and multiplying by erases every prompt position. Whether to include the assistant end-of-turn token is a data-policy choice; the key is that the mask says exactly which targets represent desired assistant behavior.
def softmax(logits, axis=-1):
"""Stable softmax."""
shifted = logits - np.max(logits, axis=axis, keepdims=True)
exp = np.exp(shifted)
return exp / np.sum(exp, axis=axis, keepdims=True)
def masked_cross_entropy(logits, targets, train_mask):
"""Mean next-token cross-entropy over masked positions, with gradient."""
logits = np.asarray(logits)
targets = np.asarray(targets, dtype=np.int64)
mask = np.asarray(train_mask, dtype=bool)
if not np.any(mask):
raise ValueError("at least one position must be trainable")
probabilities = softmax(logits, axis=-1)
rows = np.arange(targets.shape[0])
losses = -np.log(probabilities[rows, targets])
normalizer = np.sum(mask)
loss = np.sum(np.where(mask, losses, 0.0)) / normalizer
grad_logits = probabilities.copy()
grad_logits[rows, targets] -= 1.0
grad_logits *= mask[:, None] / normalizer
return loss, grad_logits
33.2 Packing without accidental examples
Short conversations waste memory if each occupies a full context window. Sequence packing concatenates multiple documents into one fixed-length row, then runs the same causal model on the row. The trap is a fake training pair at each seam: the last token of one document should not predict the first token of the next.
The safe pattern is to insert a boundary token and carry a document-id array beside the token array. A next-token position is trainable only when source and target belong to the same real document.
This is separate from attention masking. A model may be allowed to attend across packed examples for efficiency, or it may receive a block-diagonal attention mask; either way, the loss must not reward cross-document predictions. The minimal invariant is that every target with a boundary or padding document id has zero loss weight.
def packed_lm_arrays(packed_tokens, packed_doc_ids):
"""Next-token arrays whose mask forbids document-boundary targets."""
inputs = packed_tokens[:, :-1]
targets = packed_tokens[:, 1:]
same_document = packed_doc_ids[:, :-1] == packed_doc_ids[:, 1:]
train_mask = same_document & (packed_doc_ids[:, 1:] >= 0)
return inputs, targets, train_mask
This keeps the compute benefit of packing while preserving the data distribution. Padding and boundary targets get document id , so their loss mask is false. A long document may still be split across rows; then the missing transition is dropped rather than replaced by a false cross-document transition.
Packing also makes validation stricter. If two documents are accidentally joined without a boundary, the model sees a fluent but bogus continuation and the loss says it is correct. The boundary id and the document-id mask give the tests something concrete to assert: no true mask entry is allowed where adjacent ids differ.
33.3 Low-rank adaptation
Full fine-tuning updates every matrix . LoRA keeps frozen and trains a low-rank update [hu2021lora]:
The code initializes , so and the adapted layer has identical outputs at step 0. is random; once moves, both factors can cooperate. After training, the update is merged into for inference, so serving has no extra matrix multiply.
The rank controls the adapter’s capacity. Rank 1 can only add an outer product; larger ranks add a sum of such products. The scale keeps update magnitudes comparable when changes, so changing the rank does not automatically change the first useful step size. At initialization because , but is generally nonzero, so the adapter starts moving immediately without changing the initial model output.
def init_lora(d_in, d_out, rank, rng, scale=0.01):
"""Return A and zero B so the LoRA layer matches the frozen base at init."""
a = rng.normal(0.0, scale, size=(rank, d_out)).astype(np.float32)
b = np.zeros((d_in, rank), dtype=np.float32)
return a, b
def lora_output(x, w, a, b, alpha):
"""Compute X @ (W + (alpha / rank) * B @ A)."""
rank = a.shape[0]
adapted = w + (alpha / rank) * (b @ a)
return x @ adapted
def merge_lora(w, a, b, alpha):
"""Fold the trained low-rank update into W for inference."""
return w + (alpha / a.shape[0]) * (b @ a)
Let be the gradient with respect to the effective update . From ,
def lora_gradients(x, a, b, alpha, grad_y):
"""Backpropagate through the LoRA update for a loss on Y."""
scale = alpha / a.shape[0]
grad_update = x.T @ grad_y
grad_a = scale * b.T @ grad_update
grad_b = scale * grad_update @ a.T
return grad_a, grad_b
def lora_mse_loss_and_grads(x, w, a, b, alpha, target):
"""Tiny objective used by the chapter tests."""
y = lora_output(x, w, a, b, alpha)
diff = y - target
loss = 0.5 * np.mean(diff * diff)
grad_y = diff / diff.size
grad_a, grad_b = lora_gradients(x, a, b, alpha, grad_y)
return loss, grad_a, grad_b
def lora_parameter_savings(d_in, d_out, rank):
base = d_in * d_out
trainable = rank * (d_in + d_out)
return base, trainable, base / trainable
The savings are immediate. A matrix has 16,777,216 weights; rank-8 LoRA trains weights, a 256x reduction, computed and asserted by the tests. The base weights still dominate memory during training.
The gradient formulas are also a useful shape check. has the same shape as because is . has the same shape as because is . A finite-difference check on both factors catches transposes, missing scales, and accidental updates to the frozen matrix.
33.4 QLoRA in one paragraph
QLoRA stores the frozen base model in 4-bit NormalFloat (NF4) quantized blocks, dequantizes them for computation, and trains LoRA adapters on top [dettmers2023qlora]. Because the quantized base is frozen, gradients and optimizer state are needed only for the adapter weights. The price is quantization error and extra dequantization machinery; the benefit is that larger base models fit in the same accelerator memory.
Conceptually, QLoRA changes storage, not the supervised objective. The same chat template, assistant-only mask, packing boundaries, and LoRA merge logic still determine what the model learns.
|
In practice
|
Instruction-tuned systems use fixed templates for system, user, assistant, and tool messages; changing the template after training changes the model’s input distribution. The assistant-only loss is standard for demonstration data because the prompt is conditioning context, not behavior to imitate. LoRA adapters are commonly trained per task or customer and then merged or selected at inference, while QLoRA is useful when memory, not arithmetic, is the bottleneck [hu2021lora] [dettmers2023qlora]. RLHF pipelines often start from an SFT model before preference optimization [ouyang2022training]. |
33.5 Teach it
The one-sentence version. SFT teaches the model which replies to write, and LoRA makes that update low-rank so only a small adapter is trained.
An analogy. The prompt is the exam question and the assistant message is the answer key. You let the student read the whole question, but you grade only the answer.
At the board.
-
Draw a chat as role-marked tokens: system, user, assistant.
-
Put mask zeros over prompt tokens and ones over assistant tokens; derive the softmax gradient with the mask multiplier.
-
Pack two short documents, then cross out the seam so no target crosses it.
-
Factor a large update matrix as ; set to start from the base model exactly.
Misconceptions to address.
-
"The model should learn to predict user prompts." Not in SFT; prompts condition the answer.
-
"Packing is just concatenation." It also needs boundary-aware loss masks.
-
"LoRA changes inference forever." It can be merged into the base matrix.
Check for understanding. Why does make a LoRA layer identical to the frozen layer at initialization even when is random?
33.6 Exercises
Explain why training and inference must use the same role markers. Then say which tokens are context and which tokens are targets in a user-assistant SFT example.
Given two tokenized documents and with boundary token 99 and padding token 0, pack them into length-4 rows. Write the next-token targets and the Boolean loss mask.
For , derive the gradients for and . Then use the chapter code to compute the parameter saving for and .
References
-
[dettmers2023qlora] T. Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs. 2023. arXiv:2305.14314
-
[hu2021lora] E. J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. 2021. arXiv:2106.09685
-
[ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155