Chapter 8
Hypothesis Testing
Confidence intervals, the bootstrap, significance tests, and how many examples a model comparison needs.
A benchmark score is an estimate, not a property of a model. In 2026, LLM systems are often separated by a few percentage points on finite evaluation sets, so sampling noise can decide which line of a leaderboard looks best. Hypothesis testing gives the small set of tools needed to attach uncertainty to those scores and to compare two models on the same examples.
8.1 Sampling distributions and standard error
Treat each evaluation example as a draw from a population of tasks. For accuracy, define when the model is correct and otherwise. The measured accuracy is the sample mean . If the examples are independent and each has success probability , then
The derivation is only the variance-of-a-mean identity from Section 6.3: Bernoulli variables have variance , and averaging independent copies divides variance by . The square root is the standard error, the typical wiggle of across repeated evaluation sets. The population is an abstraction: it might be all user questions your product will face, all programming issues in a domain, or all prompts matching the benchmark’s sampling recipe. The standard error is therefore conditional on that sampling story; it does not account for benchmark leakage, bad labels, or distribution shift. For 248 correct answers out of 400, and the estimated standard error is 0.024.
def accuracy_standard_error(correct):
"""Standard error of the mean of 0/1 correctness indicators."""
correct = np.asarray(correct, dtype=np.float64)
p_hat = correct.mean()
return np.sqrt(p_hat * (1.0 - p_hat) / correct.size)
def normal_accuracy_ci(correct, z=1.96):
"""Normal-approximation confidence interval for an accuracy."""
correct = np.asarray(correct, dtype=np.float64)
p_hat = correct.mean()
half_width = z * accuracy_standard_error(correct)
return p_hat - half_width, p_hat + half_width
8.2 Confidence intervals
A confidence interval is a procedure, not a probability statement about this one model. A 95% procedure covers the population score in 95% of repeated evaluation sets, if its assumptions hold. The central limit theorem makes a normal interval reasonable for many benchmark accuracies:
For the 248-of-400 example, this interval runs from 0.572 to 0.668. It is wide enough that a reported accuracy of 0.64 from another run is not obviously better.
The bootstrap replaces the normal approximation with resampling. Resample the evaluation rows with replacement, recompute the statistic, and take quantiles of those bootstrap statistics. It works for accuracy, median judge score, pass@k variants, or any metric that is a function of examples, as introduced by Efron [efron1979]. The bootstrap assumes the observed rows are representative enough that drawing from them mimics drawing from the population. It cannot fix a benchmark whose rows all test the wrong skill, but it can expose how much the reported score moves when the finite set is perturbed.
def bootstrap_ci(values, statistic=np.mean, draws=2000, level=0.95, rng=None):
"""Percentile bootstrap interval for any statistic of examples."""
values = np.asarray(values)
rng = np.random.default_rng(0) if rng is None else rng
n = len(values)
stats = np.empty(draws, dtype=np.float64)
for draw in range(draws):
sample = values[rng.integers(0, n, size=n)]
stats[draw] = statistic(sample)
alpha = (1.0 - level) / 2.0
return tuple(np.quantile(stats, [alpha, 1.0 - alpha]))
8.3 Significance tests and p-values
A significance test starts with a null hypothesis , chooses a statistic , and asks how extreme the observed statistic would be if were true:
For two models on the same examples, a permutation test uses the null hypothesis that the two labels are exchangeable within each row. Flip a fair coin for every row, swap the two models' correctness on heads, and recompute the accuracy gap. The p-value is the fraction of shuffled gaps at least as large as the observed one. It is not the probability that the null is true; it is the probability of data this extreme under the null. Small p-values are strongest when the test was chosen before seeing the results; after-the-fact slicing changes the game.
8.4 Paired comparisons
Do not compare two benchmark accuracies as if they came from unrelated test sets when both models answered the same prompts. Most examples are easy or hard for both models, so the paired differences have lower noise than two independent accuracies.
The paired bootstrap resamples rows and computes each time. For binary correctness, McNemar’s test goes even further: it discards rows where both models agree and keeps only discordant rows. If is the count where A is right and B is wrong, and where B is right and A is wrong, the null says each discordant row is equally likely to favor either model. The exact two-sided p-value is a binomial tail:
def paired_bootstrap_delta(a_correct, b_correct, draws=2000, level=0.95, rng=None):
"""Bootstrap interval for accuracy(B) - accuracy(A) on paired examples."""
delta = np.asarray(b_correct, dtype=np.float64) - np.asarray(a_correct,
dtype=np.float64)
rng = np.random.default_rng(0) if rng is None else rng
n = len(delta)
estimates = np.empty(draws, dtype=np.float64)
for draw in range(draws):
estimates[draw] = delta[rng.integers(0, n, size=n)].mean()
alpha = (1.0 - level) / 2.0
return delta.mean(), tuple(np.quantile(estimates, [alpha, 1.0 - alpha]))
def exact_permutation_p_value(a_correct, b_correct):
"""Two-sided sign-flip test for paired 0/1 correctness arrays."""
delta = np.asarray(b_correct, dtype=int) - np.asarray(a_correct, dtype=int)
signs = delta[delta != 0]
observed = abs(signs.sum())
if len(signs) == 0:
return 1.0
extreme = 0
for wins_for_b in range(len(signs) + 1):
total = 2 * wins_for_b - len(signs)
if abs(total) >= observed:
extreme += comb(len(signs), wins_for_b)
return extreme / (2 ** len(signs))
def mcnemar_exact_p_value(a_correct, b_correct):
"""Exact two-sided McNemar p-value for discordant paired outcomes."""
a = np.asarray(a_correct, dtype=bool)
b = np.asarray(b_correct, dtype=bool)
a_only = int(np.sum(a & ~b))
b_only = int(np.sum(~a & b))
n = a_only + b_only
if n == 0:
return 1.0
tail = sum(comb(n, k) for k in range(min(a_only, b_only) + 1)) / (2 ** n)
return min(1.0, 2.0 * tail)
In a tiny benchmark where B fixes four A errors and breaks one A success, the observed gap is 0.25 and both the exact permutation test and McNemar’s test give . That is not evidence strong enough to trust the gap. If you inspect many models, prompt variants, metrics, or slices, adjust the story for multiple comparisons; one low p-value among many tries is easy to manufacture by chance [wasserman2004].
8.5 How many examples?
Before collecting evaluations, decide the smallest difference worth detecting. Rearranging (8.1) and (8.2) gives the sample-size arithmetic for a target half-width :
The conservative choice is , where is largest. A 95% interval with half-width 0.02 then needs 2,401 examples. If you expect accuracy near 0.70, the same arithmetic needs 2,017 examples. Paired comparisons can need fewer examples when models make errors on the same rows, but the honest estimate then comes from a pilot set of per-row differences.
def accuracy_sample_size(p, half_width, z=1.96):
"""Examples needed so a normal CI has about the requested half-width."""
return int(np.ceil(z * z * p * (1.0 - p) / (half_width * half_width)))
|
In practice
|
Agent and coding benchmarks such as SWE-bench [jimenez2023swebench] and judge-based LLM comparisons such as MT-Bench and Chatbot Arena [zheng2023judging] are finite samples of a larger task population. Report intervals for the metric, not only a point estimate, and use paired tests when the same prompts are answered by both systems; the prompt is the natural blocking variable. Bootstrap rows, not tokens: the unit of evidence is usually the task, conversation, or user query. Keep a final test set untouched; repeated leaderboard peeking turns hypothesis testing into model selection. |
8.6 Teach it
The one-sentence version: a benchmark score is a noisy sample mean, so compare models by the noise of the examples, not by wishful decimals. Analogy: judging a model from a benchmark is like judging a coin from flips; more flips shrink uncertainty, but never to zero. Board steps: write correctness as 0/1 variables; derive ; draw the normal interval; for two models, replace raw scores with per-example differences. Misconceptions: 95% confidence does not mean a 95% chance this fixed interval contains the truth; a p-value is not the probability the null is true; unpaired tests waste information on paired benchmarks. Check for understanding: if two models answer exactly the same examples, what array should you bootstrap to estimate the uncertainty of their accuracy gap?
8.7 Exercises
A model answers 248 of 400 benchmark examples correctly. Compute its accuracy, standard error, and 95% normal-approximation interval. Explain what the standard error measures.
Why can a percentile bootstrap interval be used for a median judge score even when the normal accuracy interval is not appropriate? State the resampling unit for an LLM benchmark.
On 12 examples, model B fixes four errors made by model A and breaks one example that A had right. The other rows agree. Compute the accuracy gap and the exact paired p-value.
Write a function that returns the number of examples needed for a 95% normal interval with a chosen half-width. Evaluate it for worst-case accuracy and half-width 0.02.
References
-
[efron1979] B. Efron. Bootstrap methods: another look at the jackknife. Annals of Statistics 7(1), 1–26, 1979.
-
[mcnemar1947] Q. McNemar. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153–157, 1947.
-
[wasserman2004] L. Wasserman. All of Statistics: A Concise Course in Statistical Inference. Springer, 2004.
-
[jimenez2023swebench] C. E. Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? 2023. arXiv:2310.06770
-
[zheng2023judging] L. Zheng et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. 2023. arXiv:2306.05685