Statistical Inference: Hypothesis Testing
After this lesson, you will be able to:
- Set up a hypothesis test: state the 'boring explanation' (nothing special is happening) and the 'interesting explanation' (something real is going on)
- Understand p-values: the probability of a result at least this extreme IF the boring explanation were true. A tiny p-value means the data would be very surprising under the boring explanation, but it does not prove anything for certain
- Compute a test statistic by hand: (estimate - null value) / standard error, the number of standard errors between what you saw and what the boring explanation predicts
- Tell apart the two ways a test can be wrong: a false alarm (saying something works when it does not) vs. a missed discovery (saying nothing works when it actually does)
- Build and read confidence intervals: a range of values that is likely to contain the true answer, and see how they line up with p-values
Before You Start
#The Core Question
This lesson gives you a superpower: the ability to tell real patterns from lucky coincidences. Every time someone says "this new feature increased sales" or "this drug cures patients," hypothesis testing is how you know if they are right or just fooling themselves.
The idea is simple: assume the boring explanation is true (they are guessing), calculate how unlikely the observed result would be under that assumption, and if it is very unlikely, reject the boring explanation.
#Null Hypothesis vs Alternative Hypothesis
Every hypothesis test starts with two competing claims:
- Null hypothesis (H0): The "nothing interesting is happening" explanation. Your friend is guessing randomly. The drug has no effect. The new algorithm is no better than the old one.
- Alternative hypothesis (H1 or Ha): The "something real is happening" explanation. Your friend has genuine ability. The drug works. The new algorithm is better.
The null hypothesis is the default. You assume it is true unless the evidence is strong enough to reject it. This is like a courtroom: the defendant (null hypothesis) is innocent until proven guilty beyond reasonable doubt.
#The p-value: Measuring Surprise
Start with the Coke versus Pepsi test and simply count. If your friend is only guessing, each of the 10 answers is a fair 50/50 coin flip, so there are 2 × 2 × ... × 2 = 2^10 = 1,024 equally likely patterns of right and wrong answers. Of those patterns:
- 45 have exactly 8 right (the number of ways to choose which 8 of the 10 answers are right)
- 10 have exactly 9 right
- 1 has all 10 right
That is 45 + 10 + 1 = 56 patterns that do as well as your friend or better, out of 1,024 patterns. So 56 / 1,024 = 0.0547: a pure guesser does this well or better about 5.5% of the time.
#How to interpret p-values
- p < 0.05 — The result is "statistically significant" at the 5% level. If the null hypothesis were true, you would see a result this extreme less than 5% of the time. Most scientists reject H0 at this threshold.
- p < 0.01 — Stronger evidence against H0. If H0 were true, a result this extreme would show up less than 1% of the time.
- p < 0.001 — Very strong evidence. Less than 0.1% under H0.
- p = 0.30 — Not significant. If H0 were true, 30% of experiments would give a result at least this extreme. Nothing unusual here.
#The picture and the formula for Coke vs Pepsi
The bars below show the probability of each possible score for a pure guesser. The red bars (8, 9 and 10 correct) are the tail whose total area is the p-value. Run the cell and read the printed total.
So the p-value is about 0.055 (exactly 0.0547). At the traditional alpha = 0.05 threshold, you would NOT reject the null hypothesis. Your friend's performance is suggestive but not conclusive. If they got 9 out of 10 right, the p-value drops to 0.011 — now that is significant.
#One-sided or two-sided?
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself. The first line imports the binomial distribution;sf(7, 10, 0.5)is "the probability of MORE than 7 successes in 10 tries at p = 0.5", which is the same as "8 or more".
from scipy.stats import binom
binom.sf(7, 10, 0.5) # P(8 or more) = 0.0546875
binom.sf(8, 10, 0.5) # P(9 or more) = 0.0107421875#The Test Statistic: How Many Standard Errors Away?
#Coke vs Pepsi, by hand
- The estimate is your friend's hit rate: 8 / 10 = 0.80.
- The null value is what a guesser scores: 0.50.
- The standard error under the null is sqrt(0.5 × 0.5 / 10) = sqrt(0.025) = 0.158.
- The test statistic is (0.80 − 0.50) / 0.158 = 1.90. Your friend sits 1.9 standard errors above a guesser.
- The p-value is the probability of landing that far out or farther if the null were true. For 10 trials we already have the exact answer from counting: 56 / 1,024 = 0.0547, and
binom.sf(7, 10, 0.5)printed the same number. - Compare with alpha = 0.05: 0.0547 is above it, so we fail to reject H0.
One honest caution. The bell-curve (normal) table gives 0.029 for a z of 1.90, which does not match 0.0547. With only 10 coin flips the true bars are chunky and the smooth curve is a poor stand-in, so for tiny samples use the exact count. With big samples the curve is accurate, as the next example shows.
#Two groups, by hand: is 91.2% really better than 90.8%?
The Sampling lesson ended with two models on test sets of n = 2,000 examples each: Model A at 90.8% accuracy and Model B at 91.2%. It promised a proper test. Here it is, using the same recipe. The null hypothesis is "the two models are equally accurate", so the null value for the difference is 0.
- Estimate: 0.912 − 0.908 = 0.004 (0.4 percentage points).
- Null value: 0.
- SE of Model A: sqrt(0.908 × 0.092 / 2000) = 0.00646. SE of Model B: sqrt(0.912 × 0.088 / 2000) = 0.00633.
- SE of the difference: sqrt(0.00646² + 0.00633²) = 0.00905.
- Test statistic: 0.004 / 0.00905 = 0.44. The gap is less than half a standard error from zero.
- Two-sided p-value: 2 × P(Z ≥ 0.44) = 0.66. If the models were equally accurate, 66% of such comparisons would show a gap at least this large.
- Compare with alpha = 0.05: far above it, so we fail to reject H0.
The cell below runs the same six lines in Python and adds a 95% confidence interval for the difference (explained in the next sections).
The honest reading: a 0.4-point gap on test sets this size is well inside the noise, and the 95% interval for the difference runs from about −1.4 to +2.2 points, so even the sign of the true difference is unknown. A claim such as "p = 0.03" for this gap would need roughly 48,000 test examples per model, not 2,000 (the next lesson shows how to compute that). If both models were scored on the very same 2,000 examples, a paired test (which compares them example by example) can be sharper than this unpaired recipe, because the models tend to make the same mistakes; the recipe above is the safe baseline.
#Small samples: the t distribution and degrees of freedom
#Effect size: how big, not just whether
A new drug works in 52% of patients vs 48% for placebo, with 100 patients in each group. Is this likely to be statistically significant?
The picture below draws a normal test statistic (a z value), not the Coke vs Pepsi counts. Drag the statistic along the axis and watch the shaded tail area, which is the p-value. Then change alpha and the one-sided / two-sided switch, and notice that the verdict flips exactly when the statistic crosses the cutoff.
A study reports p = 0.03 for a new drug. Which interpretation is correct?
#Type I and Type II Errors
No test is perfect. There are two ways to be wrong:
#Type I Error (False Positive): Convicting an Innocent Person
#Type II Error (False Negative): Letting a Guilty Person Go Free
#The Error Matrix
| H0 is actually TRUE | H0 is actually FALSE | |
|---|---|---|
| Reject H0 | Type I Error (alpha) | Correct (Power) |
| Fail to reject H0 | Correct | Type II Error (beta) |
Every hypothesis test falls into one of these four cells. You never know which one you are in — you can only control the probabilities.
#Real-World Consequences
- Medical testing: A Type I error (false positive) on a cancer screening means unnecessary biopsies and anxiety. A Type II error (false negative) means a missed cancer diagnosis. Medical tests are tuned to minimize Type II errors (high sensitivity) even at the cost of more false positives.
- Spam filters: A Type I error sends a real email to spam (annoying). A Type II error lets spam into your inbox (also annoying). Email providers balance both based on what users hate more.
- ML model comparison: A Type I error means you deploy a "better" model that is actually the same. A Type II error means you stick with the old model when the new one is genuinely better.
An A/B test reports p = 0.20 for 'B is better than A.' Which is the right conclusion?
#Run a real two-sample t-test
scipy.stats.ttest_ind to compute the t-statistic, p-value, and a 95% confidence interval for the difference of means — exactly the workflow you would use at any growth team. It also computes the t-statistic by hand, (estimate − null value) / standard error, so you can see that scipy's number is the same recipe from the sections above.#What a p-value does when nothing is going on
When the null hypothesis is true, the p-value is uniformly distributed between 0 and 1. Every value is equally likely, so p < 0.05 happens 5% of the time and p < 0.5 happens half the time. That is exactly what alpha = 0.05 promises. (This is exact for continuous measurements like session minutes; for a handful of coin flips the possible p-values come in coarse steps.)
Here is the five-line proof by simulation: compare two groups that come from the same distribution, 5,000 times, and histogram the p-values.
You should see roughly 500 p-values in each of the ten bins and about 4.7% below 0.05 (the printed values are 486, 504, 485, 509, 517, 491, 539, 487, 496, 486 and 0.0466). Each ticket in the lottery is equally likely, and 5% of them are "significant" by chance alone.
#Confidence Intervals: A Range That Probably Contains the Truth
Start with one number. Model B scored 91.2% on n = 2,000 examples. Its standard error is sqrt(0.912 × 0.088 / 2000) = 0.00633. The margin of error is 1.96 × 0.00633 = 0.0124, about 1.24 percentage points, so the 95% confidence interval is 91.2% ± 1.24 points, from 90.0% to 92.4%. This matches the range in the Sampling lesson. Now the general formula:
#What a 95% Confidence Interval Actually Means
A 95% confidence interval does NOT mean "there is a 95% probability that the true value is in this range." That is the most common misinterpretation in all of statistics.
The method catches the true value 95% of the time in the long run. This one interval either does or does not.
If you repeated this experiment many times and computed a CI each time, 95% of those intervals would contain the true value. Any single interval either contains the truth or it does not — you just do not know which. Each horizontal bar in the picture below is one repeated experiment; run it and count how many bars miss the true value.
#Confidence intervals and p-values are two views of one test
A 95% confidence interval excludes the null value exactly when the two-sided p-value is below 0.05.
Check it with the model comparison from before. For 90.8% versus 91.2%, the 95% interval for the difference is −1.4 to +2.2 points. It contains 0, and the two-sided p-value is 0.66. For a second pair, 90.8% versus 93.0% (still n = 2,000 each), the difference is 2.2 points, the standard error is 0.0086, the test statistic is 2.55 and the two-sided p-value is 0.011. The interval is +0.5 to +3.9 points, which excludes 0, and 0.011 is below 0.05. Run the cell to see both side by side.
The two answers always agree, because the interval and the p-value are built from the same difference and the same standard error. The interval is the richer of the two: it says not only "different from zero or not" but also how large the effect plausibly is.
#Connection to Machine Learning
#A/B Testing
Every tech company runs A/B tests: show version A to half the users and version B to the other half. Hypothesis testing determines whether the difference in click rates, engagement, or revenue is real or noise. If p < 0.05, ship it. If not, the change probably is not worth the complexity.
#Model Comparison
Is Model B really better than Model A? If Model B gets 91.2% accuracy vs Model A's 90.8% on 2,000 test examples each, the test above gives z = 0.44 and p = 0.66: the 0.4-point gap is well inside the noise, and nobody can tell from this data that B is better. Without the test, you might deploy a "better" model that just got lucky on the test set. Cross-validation combined with paired t-tests or bootstrap tests gives you sharper answers when the models are scored on the same examples.
#Training vs Chance
When a model achieves 65% accuracy on a binary classification problem, is it actually learning, or could a coin flip do the same? The null hypothesis is "the model is no better than random chance" (50% for binary). A hypothesis test tells you whether the model has learned real patterns. This is especially important for imbalanced datasets where a model predicting "not fraud" 99% of the time gets 99% accuracy without learning anything.
#Feature Importance
Is this feature actually useful for prediction? Statistical tests (like t-tests and chi-squared tests) help determine whether a feature has a real relationship with the target variable or just a spurious correlation. Dropping noise features improves model generalization.
#Try It Yourself
Tests · p_value_upper(8, 10) should be approximately 0.0547 and p_value_upper(9, 10) approximately 0.0107, so the n = 10 cutoff is 9 heads. For n = 10 and a true p of 0.7 the exact power is about 0.1493 (the simulated rate with seed 42 is 0.147). For n = 100 the cutoff is 59 heads, the exact power is about 0.9928 and the simulated rate is 0.993.
The printed values are the real ones. At n = 10 the rule is "reject H0 when you see 9 or more heads" (8 or more has a p-value of 0.0547, just above 0.05, while 9 or more has 0.0107). A coin that lands heads 70% of the time gives 9 or more heads only 14.9% of the time (simulated: 0.147), so the test misses the bias about 85% of the time. That is a Type II error rate of 85%, or low power. A two-sided version of the same rule would reject on 9 or more heads or 1 or fewer, and its power is also about 15%. At n = 100 the cutoff becomes 59 heads and the power is 0.9928, so the test almost always catches the bias. (Be careful with older rules of thumb that say "30 to 40%" for n = 10: that figure belongs to a looser cutoff of 8 heads, which has power 0.383 but also a false-alarm rate of 0.0547, above the alpha of 0.05 we chose.)
Key insight: p-values depend on BOTH effect size and sample size. A real but small effect needs lots of data to detect, and a large effect can be detected with very little.
⚡ Playground: Probability Explorer → — explore sampling distributions and watch the Central Limit Theorem emerge.
#Key Takeaways
- Hypothesis testing asks "could this be random noise?" — You assume the null hypothesis (no effect) is true and calculate the probability of seeing your data under that assumption; if the probability is very low, you reject the null hypothesis
- The p-value is NOT the probability that the null hypothesis is true — It is the probability of seeing data this extreme IF the null hypothesis were true; a subtle but critical distinction
- The test statistic (estimate − null value) / standard error counts how many standard errors your result sits from the null; for two groups the standard error of the difference is the square root of the sum of the two squared standard errors
- Under a true null the p-value is uniform between 0 and 1, and a 95% confidence interval excludes the null value exactly when the two-sided p-value is below 0.05
- Type I errors (false positives) and Type II errors (false negatives) are in tension — You can only reduce both by collecting more data; the significance level alpha controls the false positive rate; statistical power (1 - beta) controls the detection rate
- Confidence intervals complement hypothesis tests — Instead of just "is there an effect?", they tell you "how big is the effect?"; wider intervals mean less certainty, narrower intervals mean more data or less variability
- In ML, hypothesis testing powers A/B tests, model comparison, and feature selection — Every time you ask "is this improvement real?" you are doing statistical inference
#Quick Check
Your friend gets 8 out of 10 Coke vs Pepsi identifications correct. The one-sided p-value is 0.0547. At the alpha = 0.05 significance level, what do you conclude?