What’s one thing you learned? What’s still confusing?
Statistical Inference: Hypothesis Testing
Understand hypothesis testing, p-values, confidence intervals, and the errors that lurk in every statistical decision.
Power & Sample Size
A test that cannot detect a real effect is a coin flip with extra steps. This lesson teaches you to size an experiment before you run it: power, the four levers, sample size for an A/B test, and power curves.
Multiple Testing, Peeking & P-Hacking
A test run twenty ways will find something that is not there, and a dashboard checked every morning will eventually turn green by luck. This lesson teaches you to count how many questions you asked and to pay for each one.
Interactive Labs for This Track
Loss Landscape
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Vectors & Matrix Operations
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Probability Distributions
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Ask questions, share insights
Where Part 1 left off: in Sampling, Standard Error & the Central Limit Theorem you saw that every statistic wobbles from sample to sample, and that the standard error, sigma over the square root of n, measures the wobble of a mean. That formula comes from Part 1, and it exists for only a few statistics. This lesson asks what to do when there is no formula.
Compute whatever you care about on that one resample: the mean, the median, a correlation, a model's accuracy. This single number is one bootstrap replicate.
Do steps 1 and 2 about 1,000 to 10,000 times. You now hold thousands of replicates of the statistic. Their histogram is the bootstrap's picture of the sampling distribution.
Original sample: the sum is 160, so the mean is 16.0. Sorted they read 9, 11, 12, 13, 14, 15, 16, 18, 22, 30, so the median is the average of the 5th and 6th values, (14 + 15) / 2 = 14.5. Note the 30 pulling the mean above the median.
Now draw three bootstrap resamples, each of 10 values with replacement (shown sorted, so repeats are easy to see).
| Resample | Values (sorted) | Sum | Mean | Median |
|---|---|---|---|---|
| 1 | 9, 12, 12, 12, 15, 15, 16, 16, 16, 18 | 141 | 14.1 | (15 + 15) / 2 = 15.0 |
| 2 | 9, 11, 11, 12, 14, 14, 15, 15, 22, 30 | 153 | 15.3 | (14 + 14) / 2 = 14.0 |
| 3 | 11, 14, 14, 14, 15, 16, 18, 18, 22, 30 | 172 | 17.2 | (15 + 16) / 2 = 15.5 |
Notice what happened. Resample 1 never drew the 30 or the 22, so its mean fell to 14.1. Resample 3 drew the 30 once and 22 once but missed 9, 12 and 13, and its mean rose to 17.2. Each resample is a plausible alternative dataset, and the statistic moves between them. Three replicates are far too few to read a range from, so a computer does the same thing 10,000 times. Sort the 10,000 means and cut off the lowest 250 and highest 250. The result:
You also get a bootstrap standard error: 1.82 for the mean, close to the formula s over the square root of n, which is 6.15 divided by 3.16, or 1.94. The two agree because the formula is exactly right for the mean, up to one detail. The bootstrap treats your ten values as the whole population, and a population's standard deviation divides by n, giving 5.83 instead of 6.15. Then 5.83 divided by 3.16 is 1.84, and the bootstrap's 1.82 matches that within its own noise. The gap between 1.82 and 1.94 is the n versus n - 1 difference, not an error. For the median, the bootstrap gives you 1.75 where no tidy formula exists.
choice with replacement, and two percentile calls.Look at the median histogram: it is lumpy, not smooth. With only ten values, a resample median can only be one of a handful of numbers (an observed value or the midpoint of two). That lumpiness is an honest warning that ten observations carry limited information about a median. Try replacing the ten values with 100 of your own invention and the histogram smooths out.
The StreamBox demo from Part 1 has a second tab for this. Open the "Bootstrap" tab, resample one sample of StreamBox users 1,000 times for the mean, then switch to the median and read its interval the same way.
In a bootstrap resample of your n observations, what happens?
The bootstrap is powerful and easy to over-trust. It rests on one assumption: that your sample is a decent miniature of the population. Three situations break it.
The cell below shows the third failure in numbers. It builds a series where each value leans on the one before (phi = 0.8, like a slowly drifting sensor), measures how much its mean really wobbles across 2,000 fresh series, and compares with what the ordinary bootstrap claims from a single series.
Check it: this series has a standard deviation of 1.67, so the standard error of its mean is about 1.67 divided by the square root of 22, which is 0.35, matching the 0.348 measured above, while the independent-data formula 1.67 over the square root of 200 gives only 0.12, matching what the ordinary bootstrap reported.
The same logic holds for repeated measurements of the same user and for neighboring pixels. More dependent data helps less than the headline n suggests, and a larger sample still cannot remove bias.
Part 1 showed that accuracy is the mean of a 0/1 column, so its standard error is sqrt(p(1 - p) / n): 91% on 2,000 test examples is 91% give or take 1.25 points, and on 100 examples it is give or take 5.6 points. The bootstrap reaches the same answer without the formula. Resample the 2,000 right-or-wrong outcomes with replacement 5,000 times and the standard error comes out 0.0063 with a range of 0.897 to 0.922. Do it for 100 outcomes and the range is 0.850 to 0.960. Accuracy has a formula, but AUC, F1 and the gap between two models do not, and for those the bootstrap is the error bar.
On n = 2,000 examples, a model at 90.8% has SE 0.0065 and a model at 91.2% has SE 0.0063, so each carries about 1.3 points of 95% half-width. The two scores differ by only 0.4 points, much less than that. Do not settle it by checking whether two separate ranges overlap: that is the wrong test. Whether a 0.4 point gap is real depends on the standard error of the DIFFERENCE between the two scores, and the next lesson builds exactly that test. For now, the lesson is that two bare point estimates, with no error bar at all, cannot support a ranking.
Bagging stands for bootstrap aggregating. Train many models, each on a different bootstrap resample of the training set, then average their predictions. A single flexible model such as a deep decision tree is unstable: a small change in the data changes its answer a lot. Averaging many such models cancels that noise, exactly as averaging n observations cancels noise in a mean. A random forest is bagging plus random feature choices at each split.
Bootstrap a 91% test accuracy and compare the widths with the formula from Part 1. The next step bootstraps the mean and the median of the ten job times, and the stretch step counts how many independent observations a correlated series is worth.
Tests · Verify the bootstrap width is about 2.5 points at n = 2000 and about 11 points at n = 100, that the bootstrap standard error of the mean of the ten job times is near 1.84 (1.83 with this seed) and of the median near 1.76, and that 200 correlated values at phi = 0.8 have an effective sample size of 22.2.
Running the solution, the bootstrap on 2,000 outcomes gives a standard error of 0.0063 and a range of 0.897 to 0.922, width 2.5 points, twice the 1.25 point half-width of the formula. On 100 outcomes it gives 0.0285 and a range of 0.850 to 0.960, width 11.0 points, about 4.4 times wider for 20 times less data, as the square root law says (the square root of 20 is 4.5). On the ten job times the bootstrap standard error of the mean is 1.83, beside the 1.84 of the n divisor and the 1.94 of the n - 1 divisor, and the median's is 1.76, where no tidy formula exists. The effective sample size of 200 values is 200.0 at phi = 0, 66.7 at 0.5, 22.2 at 0.8 and 5.1 at 0.95: the stronger the lean on the previous value, the fewer independent observations you really have.
You have 8 observations and compute their median. You run a bootstrap and the 95% range is narrow. Name two reasons to be cautious before reporting that range as the true uncertainty.
A k-fold cross-validation score is an average of k scores from k different train and validation splits. Those k numbers differ for the same reason sample means differ. Report the mean and its spread across folds, not the mean alone, and remember the folds share most of their training data, so they are not independent draws and the naive formula standard deviation over square root of k understates the true uncertainty.