Chapter 29
Scaling Laws & Pretraining Recipes
Compute budgets, Chinchilla, muP, data curation, multi-token prediction, and model case studies.
Scaling laws are the budget math of pretraining. They do not tell you what architecture to invent, but they do tell you when a run is too small, too short, or wasting compute on the wrong axis. A modern recipe combines these laws with data curation, stable hyperparameter transfer, and late-stage schedule choices.
29.1 Training compute and simple power laws
For a dense decoder Transformer, a useful training-compute estimate is
Here is the number of non-embedding parameters and is the number of training tokens. The derivation is the usual accounting rule: the forward pass costs about FLOPs, and backpropagation through activations and weights costs about more. The constant is approximate, but the product form is what matters: doubling parameters or tokens doubles compute if the other axis is fixed.
This estimate deliberately ignores tokenizer details, sequence packing, attention pattern, and hardware utilization. Those details decide wall-clock time, but they obscure the first-order budget trade. If is fixed, then is fixed: spending more compute on parameters means spending less on tokens. Scaling laws are useful because they put a loss model on top of that trade instead of leaving it to guesswork.
def transformer_training_compute(parameters, tokens):
"""Approximate training FLOPs for a dense Transformer."""
forward = 2 * parameters * tokens
backward = 4 * parameters * tokens
return forward + backward, forward, backward
A basic scaling law says that a loss gap follows a power law, for example . Taking logs makes it a line:
So the smallest fitting code is just linear regression in log space. Kaplan et al. measured smooth power laws for language-model loss and argued that compute-optimal training should grow model size faster than data size [kaplan2020scaling]. Chinchilla revisited the allocation and found that many large models were undertrained on tokens; the practical conclusion shifted toward growing parameters and data together [hoffmann2022training].
The log-space fit also shows the limitation of a single-axis law. If you train several model sizes for several token budgets, the loss is not a function of alone or alone. A small model can be saturated by more data, and a large model can be starved by too little data. The next fit makes both failure modes explicit.
def fit_power_law(x, y):
"""Fit y = coefficient * x ** (-exponent) in log space."""
x = np.asarray(x, dtype=np.float64)
y = np.asarray(y, dtype=np.float64)
slope, intercept = np.polyfit(np.log(x), np.log(y), deg=1)
return float(np.exp(intercept)), float(-slope)
29.2 Chinchilla’s parametric loss
Hoffmann et al.'s Approach 3 fits the loss as
The published fit is , , , , and ; these values are checked in the tests against the paper’s Approach 3 table before being used here [hoffmann2022training]. The first term is the irreducible floor, the second is the penalty for too few parameters, and the third is the penalty for too few tokens.
The formula is not a promise that every dataset follows the same constants. It is a local model of a family of Transformer runs under a particular training recipe. Its value is that it turns an expensive question—which pair should we try?--into a differentiable constrained optimization problem. The result is a frontier: for each compute budget, points away from the frontier waste loss on either an overlarge model with too little data or an undersized model trained for too long.
def chinchilla_loss(parameters, tokens, constants=CHINCHILLA):
"""Approach-3 Chinchilla loss fit."""
return (constants["E"]
+ constants["A"] / parameters ** constants["alpha"]
+ constants["B"] / tokens ** constants["beta"])
def chinchilla_optimal_allocation(compute, constants=CHINCHILLA):
"""Minimize the parametric loss under compute ~= 6ND."""
a, b = constants["alpha"], constants["beta"]
A, B = constants["A"], constants["B"]
product = compute / 6
parameters = ((a * A) / (b * B) * product ** b) ** (1 / (a + b))
tokens = product / parameters
return parameters, tokens
def chinchilla_rule_tokens(parameters):
"""The common Chinchilla rule of thumb: about 20 tokens per parameter."""
return 20 * parameters
To allocate a compute budget, write so the constraint is . Substitute into (29.3) and differentiate. The optimum satisfies
Solving those two equations gives and . Because the exponents are close, the frontier grows parameters and tokens at nearly the same rate over the fitted range. The rule of thumb that survived into practice is about training tokens per parameter: a B-parameter model would get about T tokens. The exact parametric optimum is compute-dependent, so treat the rule as a planning heuristic, not a law of nature.
The stationarity condition has a useful interpretation. The left side is the scaled parameter penalty remaining in the loss, and the right side is the scaled data penalty. If the data side is larger, another token is more valuable than another parameter; if the parameter side is larger, the model is too small for the available data. Compute-optimal training equalizes those marginal returns under the budget.
29.3 Recipe details beyond size
Scaling laws assume the data distribution and optimizer recipe are fixed; real pretraining changes both. μP, or maximal update parameterization, chooses width scalings so that learning rates and other hyperparameters tuned on small models transfer to wider models [yang2022tensor]. In practice, μP is a way to spend fewer expensive large-model trials: tune a proxy, then scale width without retuning every knob.
Data quality changes the effective token count. FineWeb focuses on filtering and deduplicating web text into higher-quality pretraining data [penedo2024fineweb], while DataComp-LM studies how dataset construction choices affect language-model training sets [li2024datacomplm]. Cleaner data can move a run down the loss curve without changing or raw .
This is why recipe papers report filters, deduplication, document quality classifiers, and mixture weights instead of only token counts. Huge volumes of boilerplate are not the same training signal as diverse, well-filtered text. Scaling laws remain useful, but the effective is a property of the data pipeline, not merely a byte counter.
Schedules also matter after the main law picks a scale. A WSD schedule warms up, holds a stable learning rate, then decays for an annealing phase; the scaling book treats these schedule choices as part of the compute recipe rather than decoration [scalingbook2025]. Mid-training changes the mixture, context length, or objective after broad pretraining, and annealing spends the last compute on a lower learning rate or cleaner mix.
WSD is popular because it separates jobs that conflict in one smooth curve. Warmup avoids early optimizer shocks, the stable region performs most high-throughput learning, and the decay region trades speed for a cleaner final point. Mid-training often sits near the boundary between the stable and decay phases: the model has broad competence, so changing the distribution can specialize it without paying for a full restart.
def warmup_stable_decay(step, warmup, stable, total):
"""A simple WSD learning-rate multiplier."""
if step < warmup:
return step / warmup
if step < stable:
return 1.0
progress = (step - stable) / max(total - stable, 1)
return 0.5 * (1 + np.cos(np.pi * min(progress, 1.0)))
DeepSeek-V3 adds a multi-token prediction objective during pretraining, asking the model to predict future tokens beyond the next one [deepseekai2024deepseekv3]. That kind of auxiliary objective tries to extract more learning signal per token, but it does not remove the need to budget , , and data quality together.
|
In practice
|
Use scaling laws before launching a run, not after it fails. First estimate the compute budget with . Then choose a parameter/token pair near the Chinchilla frontier, adjust for hardware and inference cost, and spend serious effort on data filtering. Finally, reserve enough budget for schedule transitions: context extension, data-mixture changes, and a decay or annealing phase can be decisive even when the headline and look right. |
29.4 Teach it
The one-sentence version. Scaling laws turn a pretraining budget into a parameter count, token count, and recipe that are unlikely to waste the run.
An analogy. You are packing for a long trip. Kaplan says bigger suitcase first; Chinchilla says do not buy a giant suitcase and forget the clothes. Data quality is packing useful clothes instead of newspaper.
At the board.
-
Derive as forward plus backward compute.
-
Fit against and read the slope as a power-law exponent.
-
Write Chinchilla’s two penalty terms and the constraint .
-
Differentiate to get .
Misconceptions to address. The -tokens-per-parameter rule is a heuristic, not a constant of physics. More raw tokens are not always better than cleaner tokens. A schedule cannot rescue a badly undertrained or badly overlarge model.
Check for understanding. If compute is fixed and you double , what must happen to , and which loss term gets worse?
29.5 Exercises
Derive from the forward term and backward term. What happens to compute if doubles and is fixed?
Starting from , derive (29.2). Explain how the slope of a line fit in log space gives the exponent.
Implement a small grid search over candidate pairs, keep only pairs with , and return the pair with the smallest Chinchilla loss. Compare it with the closed-form allocation.
References
-
[scalingbook2025] Google DeepMind. How to scale your model. Online book, 2025. https://jax-ml.github.io/scaling-book/
-
[deepseekai2024deepseekv3] DeepSeek-AI et al. DeepSeek-V3 Technical Report. 2024. arXiv:2412.19437
-
[hoffmann2022training] J. Hoffmann et al. Training Compute-Optimal Large Language Models. 2022. arXiv:2203.15556
-
[kaplan2020scaling] J. Kaplan et al. Scaling Laws for Neural Language Models. 2020. arXiv:2001.08361
-
[li2024datacomplm] J. Li et al. DataComp-LM: In search of the next generation of training sets for language models. 2024. arXiv:2406.11794
-
[penedo2024fineweb] G. Penedo et al. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. 2024. arXiv:2406.17557
-
[yang2022tensor] G. Yang et al. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer. 2022. arXiv:2203.03466