You tracked 20 metrics, one came up with a p-value of 0.03, and you shipped. Or you checked the dashboard every morning and stopped the day it turned green. Both mistakes come from the same missing habit: deciding how many questions you are allowed to ask, and how often you may look, before you look.
Part 1, Power & Sample Size, taught you to decide the experiment before you run it: power is the chance of catching a real effect, and the sample size comes from the distance between two bells, about 14,752 users per arm to see a lift from 10% to 11%. That plan assumes you look once, at the end, and ask one question. This lesson is about what happens when you do not: peeking every day, testing many metrics, and trying many analyses.
A reminder of the notation from Part 1. A proportion such as a conversion rate is written p1, p2 or pbar, and a p-value is always written out as "p-value". In the correction formulas below, p_i and p_(i) are p-values (the i-th one and the i-th smallest). In code, proportions keep names like p1 and p2 and p-values are called pvals or pv, except that a lone p is a conversion rate in the peeking cell and an array of p-values in the corrections cell.
Learning Objectives
After this lesson, you will be able to:
Explain why checking an experiment every day and stopping at the first p < 0.05 inflates false positives, and measure by how much
Control the false positive rate across many tests with Bonferroni, Holm and Benjamini-Hochberg, and apply them to ML model and hyperparameter comparisons
Name p-hacking and the garden of forking paths, and list the defences that work: pre-registration, one primary metric and replication
The sample size formula assumes you look at the data once, at the end. Real teams look every day. The problem is that each look is another chance for noise to wander across the 1.96 line, and "stop when significant" turns that wandering into a guarantee of eventually winning.
The fix is not to look less carefully, it is to understand what the false positive rate becomes. Simulate it. This is an A/A experiment: both arms get the identical experience and the same 10% conversion rate, so every significant result is a false positive. Add 500 users per arm per day for 14 days, test each day, and stop the first time p < 0.05.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The measured rate for testing once comes out near 5%, as designed. The rate for stopping at the first significant daily look comes out near 22%, more than four times higher, in an experiment where nothing is going on. The exact value depends on the number of looks and the seed, but it only ever goes up with more looks.
Why 22% and not higher? If each of the 14 looks used brand-new users, the looks would be independent tests, and the chance of at least one false win would be 1 - 0.95^14 = 51.2% (the extra cell lines show 0.50 when simulated). The daily looks are not independent. Day 14 contains all of day 13's users plus 500 more, so the z-scores of neighbouring days are almost the same number: the correlation between day 13 and day 14 is about 0.96, and between day 1 and day 14 about 0.27. When one look has already crossed the line, the next one usually has too, so the chances overlap instead of adding up, and the damage lands well below the independent-tests bound. It is still more than four times the 5% you thought you had.
A Bayesian analysis (the next lesson) does not make this problem vanish. The posterior itself stays a valid summary of the data you have, but a decision rule such as "ship the first day the probability that B is better exceeds 95%" still has an error rate that grows with the number of looks, and it depends on the prior. Whatever method you choose, simulate the stopping rule on data where nothing is going on.
Quick check
You planned 30,000 users and the p-value was 0.07 at 20,000. A teammate says 'we are close, run it to 40,000 and see.' What is the sound response?
Peeking is one way to ask many questions of the same data. The other is asking many questions at once: 20 metrics, 12 user segments, 6 variants. If each test uses alpha = 0.05 and all the nulls are true, each has a 5% chance of a false alarm, and the chance that at least one fires grows quickly. This is the family-wise error rate (FWER).
Why exactly 5%? When the null is true, the p-value is uniformly distributed between 0 and 1: it is as likely to land in 0 to 0.05 as in 0.50 to 0.55. So P(p-value below 0.05) = 5% for every test whose null is true, whatever the data look like.
FWER=1−(1−α)m
Twenty metrics, 64% chance of at least one false positive, when nothing is happening. The formula assumes the tests are independent. If they are positively correlated (two revenue metrics, or the daily looks above that share data), the true chance is lower than this, so treat 1 - 0.95^m as the independent-tests figure. Bonferroni's fix below does not need independence. Check the arithmetic and then check it by simulation:
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The simulation matches the formula to two decimals. Each simulated p-value is a draw from the uniform distribution, which is exactly what a true null produces. By 50 tests, a false "win" is almost certain; by 100, it is virtually guaranteed.
There are two different promises you can ask a correction to make, and the choice matters.
Control the FWER: keep the probability of any false positive at or below alpha. Right when a single false claim is costly (a medical claim, a launch decision).
Control the FDR (false discovery rate): keep the expected fraction of your discoveries that are false at or below q. Right when you are screening many candidates and can tolerate a small share of duds, as in feature or gene screening.
Bonferroni is the simple FWER method: test each of the m hypotheses at alpha / m. It is safe but conservative. Holm is uniformly more powerful with the same guarantee: sort the p-values from smallest, compare the i-th to alpha / (m - i + 1), and stop at the first one that fails. Benjamini-Hochberg (BH) controls FDR: sort the p-values, find the largest rank i whose p-value, written p(i), is at most (i / m) x q, and reject that one and everything smaller.
You tested 8 metrics. The raw p-values, sorted, are 0.001, 0.0065, 0.012, 0.021, 0.031, 0.044, 0.30 and 0.74. Take alpha = q = 0.05 and m = 8.
Rank i
p-value
Bonferroni cutoff (0.05/8)
Holm cutoff 0.05/(8-i+1)
BH cutoff (i/8) x 0.05
1
0.001
0.00625
0.00625
0.00625
2
0.0065
0.00625
0.00714
0.01250
3
0.012
0.00625
0.00833
0.01875
4
0.021
0.00625
0.01000
0.02500
5
0.031
0.00625
0.01250
0.03125
6
0.044
0.00625
0.01667
Uncorrected: six of the eight (p below 0.05) look like discoveries.
Bonferroni: only p = 0.001 passes 0.00625. One discovery.
Holm: 0.001 passes 0.00625; 0.0065 passes 0.00714; then 0.012 fails 0.00833 and you stop. Two discoveries. Holm kept the same FWER guarantee and found one more.
BH: walk from the bottom. Rank 8 fails (0.74 > 0.05), rank 7 fails, rank 6 fails (0.044 > 0.0375), rank 5 passes (0.031 is at most 0.03125, barely). So the largest passing rank is 5, and BH rejects ranks 1 to 5. Five discoveries, with at most about 5% of them expected to be false.
Rank 5 is the instructive case: 0.031 squeaks under its cutoff of 0.03125. Rank 6, with p = 0.044, is below 0.05 uncorrected but above its BH cutoff of 0.0375, so it is not a discovery. BH looks for the largest rank that passes, then rejects everything at or below it, even a smaller-ranked p-value that happened to miss its own cutoff.
Now code all three and run them on this example, then in a simulation with known truth.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
Read the table that prints. Uncorrected testing has about a 60% chance of at least one false discovery per experiment. Bonferroni and Holm pull the chance of any false discovery to about 5% and cost a lot of power (about 48% of the real effects found, against 85% uncorrected). BH allows a somewhat higher chance of any false discovery, but holds the average fraction of false discoveries under 5% while finding more real effects than Bonferroni, which is why it is the default when you are screening rather than making one launch decision.
Correcting alpha has a price in data, and the explorer from Power & Sample Size shows it. Set the "Metrics tested" slider to 20 and read the readout: Bonferroni's corrected alpha is 0.0025 (0.05 divided by 20), the uncorrected chance of at least one false positive is 0.6415, and the required n at the corrected alpha is 28,076 per arm, about 90% more than the 14,752 for one metric.
P-hacking is the deliberate or half-conscious search for a significant result: trying other metrics, other subgroups, other outlier rules, other tests, until something gets under 0.05, and reporting only that one. The statistician Andrew Gelman and the economist Eric Loken gave the milder version a name: the garden of forking paths. Even with no intent to cheat, the analysis you ran was chosen after seeing the data (which outliers to drop, which segment, which covariate, one-sided or two-sided). The p-value is calculated as if that path was the only one you could ever have taken, but a different dataset would have sent you down a different path. The effective number of tests m is much larger than the number you can see.
Warning signs:
A significant result in a subgroup (new users in one country, on one device) that was never in the plan.
A "primary metric" that changed between the plan and the report.
Outlier removal or transformations chosen after looking at their effect on p.
A p-value just under 0.05 reported with no mention of how many things were tried.
Defences that work: write a short pre-registration (hypothesis, primary metric, sample size, stopping rule, planned subgroups) before collecting data, designate one primary metric and treat the rest as exploratory, correct for the family of tests, and replicate surprising findings on fresh data. A subgroup finding is a hypothesis for the next experiment, not a conclusion from this one.
Train 50 models or hyperparameter settings and evaluate each on the same test set. The best observed score is a maximum of noisy estimates, so it is biased upward even if every model is identical. Choosing the winner by test score and then reporting that same test score is the multiple comparisons problem with m = 50. The fix is to select on a validation set, then report once on a fresh test set the selection never saw.
A public leaderboard that many teams probe repeatedly is a test set being reused thousands of times. Each submission is another look at the same data, so the top scores drift upward from fitting the test set's noise, not from better generalization. Practical defences are a private held-out set revealed only at the end, rounding or limiting how much feedback each submission returns, and limiting the number of submissions.
Online evaluation of a new model is an A/B test with all of the above: a power calculation for the primary metric, a fixed stopping rule, a guardrail set (secondary metrics, such as page load time or error rate, that must not get worse) corrected with Holm or BH, and a count of how many variants you ran in parallel. A "multi-armed" test with 6 variants against one control is 6 comparisons.
Check the first claim with code. All 50 models below have the same true accuracy of 80%, so any "winner" is pure luck.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The best of 50 equal models scores about 82% on a test set where every model's true accuracy is 80%: roughly 2 points of pure luck, more than twice the standard error of a single score. The same winner on fresh data falls back to 80%.
Build the second half of the "Experiment Planner" from Power & Sample Size. You will count how many chances for a false alarm a plan gives noise, correct alpha with Bonferroni and Benjamini-Hochberg, and pay for the correction in sample size. The sample size formula from Part 1 is reused as a function.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
Tests · fwer at alpha 0.05 is 0.050, 0.226, 0.642 and 0.923 for m = 1, 5, 20 and 50, with Bonferroni alpha per test 0.0500, 0.0100, 0.0025 and 0.0010; n_per_arm at 10% to 11% is 14,752 (15 days) for 1 metric, 21,951 (22 days) for 5 metrics and 28,076 (29 days) for 20 metrics; on the 8 p-values the counts are 6 uncorrected, 1 with Bonferroni and 5 with BH; in the simulation the smallest p-value is below 0.05 about 0.642 of the time and at most 0.0025 about 0.048 of the time.
When you run the solution you should see the chance of at least one false positive climb from 0.050 to 0.226, 0.642 and 0.923 as the tests grow from 1 to 5, 20 and 50, while the Bonferroni alpha per test falls from 0.0500 to 0.0010. Paying for that correction costs data: 14,752 per arm and 15 days for one metric, 21,951 and 22 days for five, and 28,076 and 29 days for twenty. Twenty metrics cost about 1.9 times the sample of one, because the smaller alpha raises the critical value z from 1.960 to 3.023. On the 8 p-values you get 6, 1 and 5 discoveries, and the simulation confirms the correction: with 20 true nulls, the smallest p-value falls below 0.05 in 0.642 of experiments, but at or below 0.05 / 20 in only 0.048.
Peeking breaks alpha — checking 14 times and stopping at the first p below 0.05 turned a nominal 5% false positive rate into about 22% in a simulation with no real effect (below the 51% bound for 14 independent tests, because the looks share data); fix the sample size or use a sequential method
Many tests means many chances — when the null is true a p-value is uniform between 0 and 1, so with 20 independent tests at alpha 0.05 the chance of at least one false positive is 1 - 0.95^20 = 0.64; Bonferroni (alpha / m) and Holm control that chance, Benjamini-Hochberg controls the false discovery fraction and finds more
BH on 8 p-values finds 5, Holm finds 2, Bonferroni finds 1 — the thresholds are (i / m) x q for BH, alpha / (m - i + 1) for Holm, and alpha / m for Bonferroni
In ML, the best of many is biased — the best of 50 identical 80%-accuracy models scored about 82% on one test set; select on validation and report on fresh data
Two arms have identical conversion rates. You test once a day for 14 days and ship the first time p < 0.05. In a simulation, about what fraction of such experiments would ship a variant that does nothing?
Next up: Bayesian Inference in Practice. You will learn to put a prior on a whole parameter, update it with data, and report a credible interval instead of a p-value.