Chapter 29

Scaling Laws & Pretraining Recipes

Compute budgets, Chinchilla, muP, data curation, multi-token prediction, and model case studies.

Scaling laws are the budget math of pretraining. They do not tell you what architecture to invent, but they do tell you when a run is too small, too short, or wasting compute on the wrong axis. A modern recipe combines these laws with data curation, stable hyperparameter transfer, and late-stage schedule choices.

29.1 Training compute and simple power laws

For a dense decoder Transformer, a useful training-compute estimate is

C≈6ND.(29.1)C \approx 6ND .\tag{29.1}

Here NN is the number of non-embedding parameters and DD is the number of training tokens. The derivation is the usual accounting rule: the forward pass costs about 2ND2ND FLOPs, and backpropagation through activations and weights costs about 4ND4ND more. The constant is approximate, but the product form is what matters: doubling parameters or tokens doubles compute if the other axis is fixed.

This estimate deliberately ignores tokenizer details, sequence packing, attention pattern, and hardware utilization. Those details decide wall-clock time, but they obscure the first-order budget trade. If CC is fixed, then NDND is fixed: spending more compute on parameters means spending less on tokens. Scaling laws are useful because they put a loss model on top of that trade instead of leaving it to guesswork.

Listing 29.1 Compute accounting
def transformer_training_compute(parameters, tokens):
    """Approximate training FLOPs for a dense Transformer."""
    forward = 2 * parameters * tokens
    backward = 4 * parameters * tokens
    return forward + backward, forward, backward

A basic scaling law says that a loss gap follows a power law, for example y=ax−αy = a x^{-\alpha}. Taking logs makes it a line:

log⁡y=log⁡a−αlog⁡x.(29.2)\log y = \log a - \alpha \log x .\tag{29.2}

So the smallest fitting code is just linear regression in log space. Kaplan et al. measured smooth power laws for language-model loss and argued that compute-optimal training should grow model size faster than data size [kaplan2020scaling]. Chinchilla revisited the allocation and found that many large models were undertrained on tokens; the practical conclusion shifted toward growing parameters and data together [hoffmann2022training].

The log-space fit also shows the limitation of a single-axis law. If you train several model sizes for several token budgets, the loss is not a function of NN alone or DD alone. A small model can be saturated by more data, and a large model can be starved by too little data. The next fit makes both failure modes explicit.

Listing 29.2 Fitting a power law in log space
def fit_power_law(x, y):
    """Fit y = coefficient * x ** (-exponent) in log space."""
    x = np.asarray(x, dtype=np.float64)
    y = np.asarray(y, dtype=np.float64)
    slope, intercept = np.polyfit(np.log(x), np.log(y), deg=1)
    return float(np.exp(intercept)), float(-slope)

29.2 Chinchilla’s parametric loss

Hoffmann et al.'s Approach 3 fits the loss as

L(N,D)=E+ANα+BDβ.(29.3)L(N,D) = E + \frac{A}{N^\alpha} + \frac{B}{D^\beta} .\tag{29.3}

The published fit is E=1.69E=1.69, A=406.4A=406.4, B=410.7B=410.7, α=0.34\alpha=0.34, and β=0.28\beta=0.28; these values are checked in the tests against the paper’s Approach 3 table before being used here [hoffmann2022training]. The first term is the irreducible floor, the second is the penalty for too few parameters, and the third is the penalty for too few tokens.

The formula is not a promise that every dataset follows the same constants. It is a local model of a family of Transformer runs under a particular training recipe. Its value is that it turns an expensive question—​which N,DN,D pair should we try?--into a differentiable constrained optimization problem. The result is a frontier: for each compute budget, points away from the frontier waste loss on either an overlarge model with too little data or an undersized model trained for too long.

Listing 29.3 Chinchilla loss and compute-optimal allocation
def chinchilla_loss(parameters, tokens, constants=CHINCHILLA):
    """Approach-3 Chinchilla loss fit."""
    return (constants["E"]
            + constants["A"] / parameters ** constants["alpha"]
            + constants["B"] / tokens ** constants["beta"])


def chinchilla_optimal_allocation(compute, constants=CHINCHILLA):
    """Minimize the parametric loss under compute ~= 6ND."""
    a, b = constants["alpha"], constants["beta"]
    A, B = constants["A"], constants["B"]
    product = compute / 6
    parameters = ((a * A) / (b * B) * product ** b) ** (1 / (a + b))
    tokens = product / parameters
    return parameters, tokens


def chinchilla_rule_tokens(parameters):
    """The common Chinchilla rule of thumb: about 20 tokens per parameter."""
    return 20 * parameters

To allocate a compute budget, write M=C/6M=C/6 so the constraint is ND=MND=M. Substitute D=M/ND=M/N into (29.3) and differentiate. The optimum satisfies

αAN−α=βBD−β,ND=C/6.(29.4)\alpha A N^{-\alpha} = \beta B D^{-\beta}, \qquad ND = C/6 .\tag{29.4}

Solving those two equations gives N∗N_* and D∗D_*. Because the exponents are close, the frontier grows parameters and tokens at nearly the same rate over the fitted range. The rule of thumb that survived into practice is about 2020 training tokens per parameter: a 7070B-parameter model would get about 1.41.4T tokens. The exact parametric optimum is compute-dependent, so treat the rule as a planning heuristic, not a law of nature.

The stationarity condition has a useful interpretation. The left side is the scaled parameter penalty remaining in the loss, and the right side is the scaled data penalty. If the data side is larger, another token is more valuable than another parameter; if the parameter side is larger, the model is too small for the available data. Compute-optimal training equalizes those marginal returns under the NDND budget.

29.3 Recipe details beyond size

Scaling laws assume the data distribution and optimizer recipe are fixed; real pretraining changes both. μP, or maximal update parameterization, chooses width scalings so that learning rates and other hyperparameters tuned on small models transfer to wider models [yang2022tensor]. In practice, μP is a way to spend fewer expensive large-model trials: tune a proxy, then scale width without retuning every knob.

Data quality changes the effective token count. FineWeb focuses on filtering and deduplicating web text into higher-quality pretraining data [penedo2024fineweb], while DataComp-LM studies how dataset construction choices affect language-model training sets [li2024datacomplm]. Cleaner data can move a run down the loss curve without changing NN or raw DD.

This is why recipe papers report filters, deduplication, document quality classifiers, and mixture weights instead of only token counts. Huge volumes of boilerplate are not the same training signal as diverse, well-filtered text. Scaling laws remain useful, but the effective DD is a property of the data pipeline, not merely a byte counter.

Schedules also matter after the main law picks a scale. A WSD schedule warms up, holds a stable learning rate, then decays for an annealing phase; the scaling book treats these schedule choices as part of the compute recipe rather than decoration [scalingbook2025]. Mid-training changes the mixture, context length, or objective after broad pretraining, and annealing spends the last compute on a lower learning rate or cleaner mix.

WSD is popular because it separates jobs that conflict in one smooth curve. Warmup avoids early optimizer shocks, the stable region performs most high-throughput learning, and the decay region trades speed for a cleaner final point. Mid-training often sits near the boundary between the stable and decay phases: the model has broad competence, so changing the distribution can specialize it without paying for a full restart.

Listing 29.4 A tiny warmup-stable-decay schedule
def warmup_stable_decay(step, warmup, stable, total):
    """A simple WSD learning-rate multiplier."""
    if step < warmup:
        return step / warmup
    if step < stable:
        return 1.0
    progress = (step - stable) / max(total - stable, 1)
    return 0.5 * (1 + np.cos(np.pi * min(progress, 1.0)))

DeepSeek-V3 adds a multi-token prediction objective during pretraining, asking the model to predict future tokens beyond the next one [deepseekai2024deepseekv3]. That kind of auxiliary objective tries to extract more learning signal per token, but it does not remove the need to budget NN, DD, and data quality together.

In practice

Use scaling laws before launching a run, not after it fails. First estimate the compute budget with 6ND6ND. Then choose a parameter/token pair near the Chinchilla frontier, adjust for hardware and inference cost, and spend serious effort on data filtering. Finally, reserve enough budget for schedule transitions: context extension, data-mixture changes, and a decay or annealing phase can be decisive even when the headline NN and DD look right.

Key equations
C≈2ND+4ND=6NDC \approx 2ND + 4ND = 6ND
log⁡y=log⁡a−αlog⁡x\log y = \log a - \alpha\log x
L(N,D)=E+A/Nα+B/DβL(N,D)=E + A/N^\alpha + B/D^\beta
αAN−α=βBD−β,ND=C/6\alpha A N^{-\alpha} = \beta B D^{-\beta}, \qquad ND=C/6
D≈20N(planning rule of thumb)D \approx 20N \quad \text{(planning rule of thumb)}

29.4 Teach it

The one-sentence version. Scaling laws turn a pretraining budget into a parameter count, token count, and recipe that are unlikely to waste the run.

An analogy. You are packing for a long trip. Kaplan says bigger suitcase first; Chinchilla says do not buy a giant suitcase and forget the clothes. Data quality is packing useful clothes instead of newspaper.

At the board.

  1. Derive C≈6NDC\approx 6ND as forward plus backward compute.

  2. Fit log⁡y\log y against log⁡x\log x and read the slope as a power-law exponent.

  3. Write Chinchilla’s two penalty terms and the constraint ND=C/6ND=C/6.

  4. Differentiate to get αAN−α=βBD−β\alpha A N^{-\alpha}=\beta B D^{-\beta}.

Misconceptions to address. The 2020-tokens-per-parameter rule is a heuristic, not a constant of physics. More raw tokens are not always better than cleaner tokens. A schedule cannot rescue a badly undertrained or badly overlarge model.

Check for understanding. If compute is fixed and you double NN, what must happen to DD, and which loss term gets worse?

29.5 Exercises

Exercise 29.1 ★ FLOP accounting

Derive C≈6NDC\approx 6ND from the 2ND2ND forward term and 4ND4ND backward term. What happens to compute if NN doubles and DD is fixed?

Exercise 29.2 ★★ Fitting a power law

Starting from y=ax−αy = ax^{-\alpha}, derive (29.2). Explain how the slope of a line fit in log space gives the exponent.

Exercise 29.3 ★★ Chinchilla allocation

Under the constraint ND=C/6ND=C/6, derive (29.4) from (29.3). Why does this condition balance the parameter and data penalty terms?

Exercise 29.4 ★★★ Grid-search a frontier

Implement a small grid search over candidate N,DN,D pairs, keep only pairs with 6ND≤C6ND\le C, and return the pair with the smallest Chinchilla loss. Compare it with the closed-form allocation.

References

  • [scalingbook2025] Google DeepMind. How to scale your model. Online book, 2025. https://jax-ml.github.io/scaling-book/

  • [deepseekai2024deepseekv3] DeepSeek-AI et al. DeepSeek-V3 Technical Report. 2024. arXiv:2412.19437

  • [hoffmann2022training] J. Hoffmann et al. Training Compute-Optimal Large Language Models. 2022. arXiv:2203.15556

  • [kaplan2020scaling] J. Kaplan et al. Scaling Laws for Neural Language Models. 2020. arXiv:2001.08361

  • [li2024datacomplm] J. Li et al. DataComp-LM: In search of the next generation of training sets for language models. 2024. arXiv:2406.11794

  • [penedo2024fineweb] G. Penedo et al. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. 2024. arXiv:2406.17557

  • [yang2022tensor] G. Yang et al. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer. 2022. arXiv:2203.03466