Power, Sample Size & Multiple Testing
After this lesson, you will be able to:
- Define statistical power as 1 minus beta, and predict how it moves when effect size, noise, sample size or alpha changes
- Separate a statistically significant result from an important one: a tiny effect with a huge sample gives a tiny p-value and still changes nothing
- Compute the sample size for an A/B test on a conversion rate from a formula, then confirm it by simulation
- Explain why checking an experiment every day and stopping at the first p < 0.05 inflates false positives, and measure by how much
- Control the false positive rate across many tests with Bonferroni, Holm and Benjamini-Hochberg, and apply them to ML model and hyperparameter comparisons
Before You Start
#Why Power Matters
#The Four Levers of Power
Power for a two-group comparison of means depends on four things, and only four. Let delta be the true difference between group means, sigma the standard deviation of one observation, n the number of observations per group, and alpha the significance level.
Start with a picture, before any formula. Imagine running the experiment thousands of times and recording the observed difference between the two group means each time. If there is no real effect, those differences pile up in a bell centred on 0 (the grey bell, the null). If the true difference is delta, they pile up in a second bell of the same width centred on delta (the blue bell, the alternative). The test rejects whenever the observed difference lands beyond a cut-off line. The red area of the grey bell beyond the line is alpha, the false alarms. The amber area of the blue bell short of the line is beta, the misses. The green area of the blue bell beyond the line is power.
Move delta, sigma and n in the picture below and watch the bells slide apart or narrow; the Baseline preset has power of about 29%.
norm.ppf, the inverse of Phi.)The distance between the bells, measured in standard errors, is the effect divided by the standard error of a difference of two group means. Each mean has standard error sigma over the square root of n, and the two variances add, giving a standard error of sigma times the square root of 2/n for the difference. Read the area to the right of the cut-off line under the blue bell:
Read the ratio inside the parentheses: effect divided by standard error. Every lever acts by moving that ratio, or by moving the critical value it has to beat.
- Larger effect (delta up): the signal is bigger, power rises.
- Less noise (sigma down): the signal is clearer, power rises.
- More data (n up): the standard error shrinks like 1 over the square root of n, power rises. Quadrupling n halves the standard error.
- Looser alpha (alpha up): the critical value drops, so you reject more easily. Power rises, but so does the false alarm rate. This lever is not free.
The first two combine into one number, Cohen's d = delta / sigma, the effect measured in units of noise. Here d = 0.5 / 2 = 0.25.
Baseline: effect 0.5, noise sigma 2, n = 64 per group, alpha 0.05, giving about 29% power. Which single change brings power to about 80%?
Run the cell and verify every lever with numbers.
Notice the asymmetry in the last two lines of the table: loosening alpha to 0.10 lifts power from 0.29 to 0.41, but you paid for it by doubling your false alarm rate. More data, or less noise per observation, is what lets you push both errors down together: you can hold alpha at 0.05 and still shrink beta.
#Effect Size vs Statistical Significance
A p-value answers "is the effect nonzero?" It does not answer "is the effect big?" Those are different questions, and sample size is what separates them.
One notation note for the rest of the lesson, because the letter p has two jobs in statistics. A proportion, such as a conversion rate, is always written p1 or p2 (or pbar for the average of the two). A p-value is always written out as "p-value" in the text. Inside the code cells, proportions keep names like p1 and p2, and p-values are called pvals or pv.
Two measures of how big an effect is:
- Absolute lift: the raw difference, such as a conversion rate going from 10.0% to 10.1%, a lift of 0.1 percentage points.
- Cohen's d: the difference divided by the standard deviation (for two groups, the two standard deviations combined into one, called the pooled standard deviation). Rough conventions are 0.2 small, 0.5 medium, 0.8 large, though what counts as important depends on cost and context, not on those labels.
Because the standard error shrinks as n grows, any nonzero effect becomes significant if you collect enough data. Here is a constructed example. Variant B converts at 10.1% versus 10.0% for control, and each arm has 3 million users.
With 3 million users per arm the lift gives z about 4.07 and a p-value of about 0.00005. The effect is real, and yet it is 0.1 percentage point, or 1,000 extra conversions per million visitors, with Cohen's d of 0.003. Whether that justifies the engineering cost of the change is a business question the p-value cannot answer. At 100,000 per arm the very same effect is nowhere near significant.
Team A reports p = 0.0004 for a new ranking model. Team B reports p = 0.08 for a different one. Which statement is safe?
#Sample Size for an A/B Test
First derive the answer from the two bells, for the simplest case where both groups have the same noise sigma. Let s = sigma x sqrt(2/n) be the standard error of the difference, the width of each bell. The cut-off line sits z_ standard errors to the right of the grey bell's centre, so it is at z_ x s. For 80% power, 80% of the blue bell must lie to the right of the line, so the line must also sit z_power standard errors to the left of the blue bell's centre, at delta minus z_power x s. Both sentences describe the same line, so:
delta = (z_ + z_power) x s
The distance between the bells has to cover two numbers: one to clear the false alarm line, and one to put 80% of the blue bell beyond it. Put s = sigma x sqrt(2/n) back in and solve for n:
Worked with the lesson's baseline numbers (sigma = 2, delta = 0.5): z_ + z_power = 1.960 + 0.842 = 2.802, squared is 7.85, and n = 2 x 4 x 7.85 / 0.25 = 251.2. Always round up, so 252 per group. That agrees with the quadrupling you saw earlier (256 per group gave 80.7%). Never trust a formula you have not simulated, so check it:
At 252 per group the simulated power is 0.803, near the 0.80 target, and at half that sample it is only about 0.50. The last line shows the squared effect at work: halving delta to 0.25 needs 1,005 per group, four times as many.
Now the A/B test itself. For a yes/no outcome there is no separate noise number sigma. A visitor converts or does not, and the variance of one visitor's outcome with rate p is p(1 - p). So sigma squared becomes pbar(1 - pbar), where pbar = (p1 + p2) / 2 is the average of the two rates, and delta becomes the lift p2 - p1:
Two conventions are hiding in that formula. The first is "pooled": we treat the two arms as one group with a single shared rate pbar and use one variance for both bells. That is exactly right for the grey bell, where there is no effect and both arms share one rate. For the blue bell the two arms really have different rates, 10% and 11%, with variances 0.09 and 0.0979. A more careful textbook version uses each arm's own variance for the power part; it gives 14,750.8, which rounds to 14,751. The two answers differ by one visitor because the rates are close, and they drift apart only when the lift is large compared with the baseline. This lesson and the planner below use the simpler pooled version.
Run the proportions version, then check it by brute-force simulation rather than trusting it.
#Power Curves
Two things to read off the right-hand chart. First, the curve starts at 0.05 when the true lift is zero: with no effect, "power" is just the false positive rate, which is alpha. Second, the 14,752-per-arm design reaches 80% only at the 1-point lift it was built for. If the true lift is half that, you will see it only about 29% of the time, and you will confidently report "no effect" most of the time you have a real one.
Now try it yourself: drag the lift and the sample size and watch both curves and the required n move, then raise the number of metrics to see Bonferroni inflate it. The StreamBox preset starts from the shared table's 22% churn rate.
#Peeking and Optional Stopping
Why 22% and not higher? If each of the 14 looks used brand-new users, the looks would be independent tests, and the chance of at least one false win would be 1 - 0.95^14 = 51.2% (the extra cell lines show 0.50 when simulated). The daily looks are not independent. Day 14 contains all of day 13's users plus 500 more, so the z-scores of neighbouring days are almost the same number: the correlation between day 13 and day 14 is about 0.96, and between day 1 and day 14 about 0.27. When one look has already crossed the line, the next one usually has too, so the chances overlap instead of adding up, and the damage lands well below the independent-tests bound. It is still more than four times the 5% you thought you had.
A Bayesian analysis (the next lesson) does not make this problem vanish. The posterior itself stays a valid summary of the data you have, but a decision rule such as "ship the first day the probability that B is better exceeds 95%" still has an error rate that grows with the number of looks, and it depends on the prior. Whatever method you choose, simulate the stopping rule on data where nothing is going on.
You planned 30,000 users and the p-value was 0.07 at 20,000. A teammate says 'we are close, run it to 40,000 and see.' What is the sound response?
#The Multiple Comparisons Problem
Why exactly 5%? When the null is true, the p-value is uniformly distributed between 0 and 1: it is as likely to land in 0 to 0.05 as in 0.50 to 0.55. So P(p-value below 0.05) = 5% for every test whose null is true, whatever the data look like.
The simulation matches the formula to two decimals. Each simulated p-value is a draw from the uniform distribution, which is exactly what a true null produces. By 50 tests, a false "win" is almost certain; by 100, it is virtually guaranteed.
#Bonferroni, Holm and Benjamini-Hochberg
There are two different promises you can ask a correction to make, and the choice matters.
- Control the FWER: keep the probability of any false positive at or below alpha. Right when a single false claim is costly (a medical claim, a launch decision).
- Control the FDR (false discovery rate): keep the expected fraction of your discoveries that are false at or below q. Right when you are screening many candidates and can tolerate a small share of duds, as in feature or gene screening.
#A hand-worked example with 8 p-values
You tested 8 metrics. The raw p-values, sorted, are 0.001, 0.0065, 0.012, 0.021, 0.031, 0.044, 0.30 and 0.74. Take alpha = q = 0.05 and m = 8.
| Rank i | p-value | Bonferroni cutoff (0.05/8) | Holm cutoff 0.05/(8-i+1) | BH cutoff (i/8) x 0.05 |
|---|---|---|---|---|
| 1 | 0.001 | 0.00625 | 0.00625 | 0.00625 |
| 2 | 0.0065 | 0.00625 | 0.00714 | 0.01250 |
| 3 | 0.012 | 0.00625 | 0.00833 | 0.01875 |
| 4 | 0.021 | 0.00625 | 0.01000 | 0.02500 |
| 5 | 0.031 | 0.00625 | 0.01250 | 0.03125 |
| 6 | 0.044 | 0.00625 | 0.01667 | 0.03750 |
| 7 | 0.30 | 0.00625 | 0.02500 | 0.04375 |
| 8 | 0.74 | 0.00625 | 0.05000 | 0.05000 |
- Uncorrected: six of the eight (p below 0.05) look like discoveries.
- Bonferroni: only p = 0.001 passes 0.00625. One discovery.
- Holm: 0.001 passes 0.00625; 0.0065 passes 0.00714; then 0.012 fails 0.00833 and you stop. Two discoveries. Holm kept the same FWER guarantee and found one more.
- BH: walk from the bottom. Rank 8 fails (0.74 > 0.05), rank 7 fails, rank 6 fails (0.044 > 0.0375), rank 5 passes (0.031 is at most 0.03125, barely). So the largest passing rank is 5, and BH rejects ranks 1 to 5. Five discoveries, with at most about 5% of them expected to be false.
Now code all three and run them on this example, then in a simulation with known truth.
#P-Hacking and the Garden of Forking Paths
Warning signs:
- A significant result in a subgroup (new users in one country, on one device) that was never in the plan.
- A "primary metric" that changed between the plan and the report.
- Outlier removal or transformations chosen after looking at their effect on p.
- A p-value just under 0.05 reported with no mention of how many things were tried.
#Connection to Machine Learning
#Many models, one test set (the winner's curse)
#Leaderboard overfitting
A public leaderboard that many teams probe repeatedly is a test set being reused thousands of times. Each submission is another look at the same data, so the top scores drift upward from fitting the test set's noise, not from better generalization. Practical defences are a private held-out set revealed only at the end, rounding or limiting how much feedback each submission returns, and limiting the number of submissions.
#How many eval examples you need
#Experiments in production
Online evaluation of a new model is an A/B test with all of the above: a power calculation for the primary metric, a fixed stopping rule, a guardrail set (secondary metrics, such as page load time or error rate, that must not get worse) corrected with Holm or BH, and a count of how many variants you ran in parallel. A "multi-armed" test with 6 variants against one control is 6 comparisons.
#Try It Yourself
Build the "Experiment Planner" from the top of this lesson. You will write the sample size formula as a function, add the Bonferroni correction for several metrics, and check the plan by simulation. The derived formula from the Sample Size section is the only maths you need.
Tests · plan_ab_test(0.10, 0.01) returns n_per_arm 14,752 and days 15; with 5 metrics alpha per test is 0.01 and n_per_arm 21,951 (22 days); with 20 metrics alpha per test is 0.0025 and n_per_arm 28,076 (29 days); baseline 0.02 with MDE 0.002 needs 80,683 per arm (81 days); the first MDE on the grid with n at most 8,000 is 1.4 points; the simulated power at n = 14,752 is near 0.80 (0.803 with seed 0).
When you run the solution you should see 14,752 per arm and 15 days for one metric, 21,951 and 22 days for five, and 28,076 and 29 days for twenty. Twenty metrics cost about 1.9 times the sample of one, because a smaller alpha raises the critical value z from 1.960 to 3.023, the alpha lever from the four levers. The rare-outcome case (2% baseline, 0.2-point lift) needs 80,683 per arm, 81 days at 1,000 users a day, since a 10% relative lift on a rare outcome is a tiny absolute lift. With 8,000 per arm the smallest MDE on the grid is 1.4 points, and the simulation of the 1-metric plan returns power 0.803.
#Key Takeaways
- Power is 1 minus beta, the chance of catching a real effect — it depends on effect size, noise, sample size and alpha; with effect 0.5, sigma 2 and n = 64 per group it is only about 29%, and doubling the effect, halving sigma, or quadrupling n each lift it to about 81%
- Significance is not importance — a 0.1-point lift in conversion with 3 million users per arm gives p about 0.00005 and Cohen's d of 0.003; report the effect size with a confidence interval, not just the p-value
- Two bells give the sample size — the distance between the bells must cover z for alpha plus z for power, so n per group = 2 x sigma^2 x (1.96 + 0.84)^2 / delta^2, which is 252 for sigma 2 and delta 0.5 and checks out in simulation
- Size the test before you run it — detecting a lift from 10% to 11% at alpha 0.05 and 80% power needs about 14,752 users per arm, confirmed by simulation (about 80% power, about 5% false positives); halving the MDE multiplies that by almost four
- Peeking breaks alpha — checking 14 times and stopping at the first p below 0.05 turned a nominal 5% false positive rate into about 22% in a simulation with no real effect (below the 51% bound for 14 independent tests, because the looks share data); fix the sample size or use a sequential method
- Many tests means many chances — when the null is true a p-value is uniform between 0 and 1, so with 20 independent tests at alpha 0.05 the chance of at least one false positive is 1 - 0.95^20 = 0.64; Bonferroni (alpha / m) and Holm control that chance, Benjamini-Hochberg controls the false discovery fraction and finds more
- BH on 8 p-values finds 5, Holm finds 2, Bonferroni finds 1 — the thresholds are (i / m) x q for BH, alpha / (m - i + 1) for Holm, and alpha / m for Bonferroni
- In ML, the best of many is biased — the best of 50 identical 80%-accuracy models scored about 82% on one test set; select on validation, report on fresh data, and detecting a 1-point accuracy gain from 90% needs about 13,500 examples per model, or roughly 3,900 with a paired test under the stated disagreement assumption
#Quick Check
An experiment has 25% power to detect the effect you care about and returns p = 0.30. What is the best reading?