Sampling, Standard Error & the Bootstrap
After this lesson, you will be able to:
- Tell a population from a sample, and a parameter (a fixed truth you want) from a statistic (a number you computed that changes with every sample)
- Explain why the sample mean wobbles from sample to sample, state the Central Limit Theorem in plain words with its conditions, and compute how much it wobbles with the standard error, s divided by the square root of n
- Never again confuse standard deviation (spread of the data) with standard error (wobble of the average), and know why 4 times the data buys only 2 times the precision
- Recognize a biased sampling method and understand why a huge biased sample is confidently wrong, while a small random one is honestly uncertain
- Run the bootstrap by hand and in code to get an honest range for a mean or a median, and know the three situations where it fails
- Put error bars on model accuracy and see bagging and cross-validation as sampling ideas
Before You Start
#The Core Question
A number computed from a sample is a guess about the population. How good a guess, and how can you tell? Everything in this lesson is a way of putting a size on that doubt, so that a result becomes "about 91%, give or take 1.3 points" (for a test set of 2,000 examples) instead of a falsely exact "91%".
#Population and Sample
- A parameter is a number describing the population, such as the true average order value, written with Greek letters: the mean is mu and the standard deviation is sigma. It is fixed, and almost always unknown.
- A statistic is a number you compute from the sample, such as the sample mean, written x-bar. It is known, and it changes every time you draw a new sample.
#The Sampling Distribution of the Mean
Here is the idea that unlocks everything. Suppose you could repeat your whole study many times, each time drawing a fresh sample of size n and computing its mean. Those means would not all be equal. If you collected them, they would form a distribution of their own, called the sampling distribution of the mean.
You cannot actually redo a study 2,000 times in real life, but a computer can, from a population we build ourselves. The cell below makes a deliberately ugly population (time on a page, strongly skewed to the right, most people leave quickly and a few stay forever), draws 2,000 samples of 30 people each, and plots the 2,000 means.
Try it! Run the cell, then changeN_SAMPLEto 5 and to 200 and watch the right-hand histogram.
Three things appear in that output, and each one matters.
- The means pile up around the true mean. The average of the 2,000 sample means is 59.80 against a true mean of 60.21. A sample mean has no tendency to miss in one direction, so it is called an unbiased estimator. The 0.41 gap is not a lean, it is simulation noise: the average of 2,000 means itself wobbles by about 11.42 divided by the square root of 2,000, which is 0.26, so a gap of 1.6 of those units is unremarkable. A different seed gives a different gap.
- The means are much less spread out than the individuals. The raw data has a standard deviation of 60.32 seconds, but the 2,000 sample means have a standard deviation of only 11.42. Averaging cancels noise.
- The histogram of means is a bell, even though the population is a lopsided ramp. That is the Central Limit Theorem, which the next section states properly.
The experiment above ran on a population we invented. The next demo runs on data you will use all through Track 1. StreamBox is a synthetic table of 200 app users (plan, device, minutes per session, sessions per week, whether they churned), and the demo draws samples from a 10,000-user population built to match it. Open the "Sampling distribution" tab, draw 100 samples at n = 30, then raise n to 300 and watch the bell narrow. The sampling-bias switch is there too, leave it off for now, because two sections from here you will turn it on and see what it does.
#The Central Limit Theorem
The simulation showed a bell appearing from a lopsided population. That is not luck, it is a theorem, and this lesson is where it is stated.
It comes with conditions, and the conditions are where it fails.
- The draws must be independent: one value must not tell you anything about the next.
- The source must have a finite variance: a spread that is a real, finite number. A source with such extreme outliers that its variance is infinite breaks the theorem, and a heavy-tailed one that merely comes close needs far more than 30 draws.
- n must not be tiny. A common rule of thumb is that n around 30 is enough for mildly skewed data, and strongly skewed data needs more. The theorem says what happens as n grows, not that any particular n is already there.
Here is where "about 95%" comes from. A Normal curve puts about 68% of its values within 1 standard deviation of its center, 95% within 2 and 99.7% within 3 (the 68-95-99.7 rule from Descriptive Statistics). The CLT says the sample mean is approximately Normal, and its standard deviation is sigma over the square root of n, our wobble unit. So about 95% of sample means land within 2 of those units of mu. In the simulation, 67.1% of the 2,000 means landed within 1 unit of 60.21, 94.9% within 2 and 99.5% within 3, with a Normal curve predicting 68, 95 and 99.7.
Now see it with your own hands: pick a source, set n, and draw batches of means.
#Standard Error and the Square Root Law
Before the formula, commit to a guess.
Your estimate from a sample of 100 has some typical error. You want to cut that typical error in half. How much data do you need in total?
Check it against the simulation above. Sigma was 60.32 and n was 30, so sigma divided by the square root of 30 is 11.01, and the simulated spread of the 2,000 means was 11.42. The formula and the experiment agree, within the noise you would expect from only 2,000 repeats.
#Sigma is unknown in practice, so use s
Why n - 1 and not n? The distances are measured from x-bar, and x-bar is the point that sits closest to your own values by construction, closer than the true mean mu would. So the squared distances come out a little too small on average, and dividing by n - 1 instead of n corrects for it. With n = 1000 the difference is invisible. With ten values it is a real 5%, and you will see exactly that gap in the bootstrap section below.
#The square root law, with numbers
Take a population with sigma = 15 and watch what each extra block of data buys.
| Sample size n | SE = 15 / sqrt(n) | Change from previous row |
|---|---|---|
| 25 | 3.00 | |
| 100 | 1.50 | 4 times the data, half the error |
| 400 | 0.75 | 4 times the data, half the error |
| 1,600 | 0.375 | 4 times the data, half the error |
| 10,000 | 0.15 |
The cell below measures this tax directly: for four sample sizes it draws 2,000 samples and compares the measured SE to the formula.
#Standard Deviation vs Standard Error
These two are confused more often than any other pair in statistics, and the confusion quietly ruins reports. They answer different questions.
The easiest way to keep them straight is to ask what happens as you collect more data.
- Standard deviation does not shrink. It is a property of the population. Measure a million people and the heights still vary by about 7 cm, because that is how people are.
- Standard error shrinks without limit. It measures your uncertainty about the average, and more data removes that uncertainty. With a million people the SE of the average height is 0.007 cm.
You measure the response time of 400 requests. The standard deviation is 80 ms. What does the standard error of the mean tell you?
#How You Sample Decides What You Learn
Estimate the average monthly spend of a company's customers. Option A is a simple random sample of 100 customers. Option B is an optional survey email that 1,000,000 customers could answer, and big spenders are much more likely to bother replying. Which average do you trust more?
#Three honest ways to sample
- Simple random sampling. Every member of the population has the same chance of being picked. It is the baseline that the standard error formula assumes.
- Stratified sampling. Split the population into groups that differ in ways that matter (age bands, countries, fraud versus not fraud), then sample randomly inside each group, in proportion to its size. You cannot accidentally miss a small group, and the estimate is usually more precise. This is what a
stratify=yargument does when you split a dataset for training. - Cluster sampling. Pick whole groups at random (a few schools, a few data centers), then measure everyone inside the chosen groups. It is cheap when visiting individuals is costly, but people inside a cluster resemble each other, so you get less information per person than a simple random sample would give.
#Three ways samples go quietly wrong
- Self-selection bias. People choose whether to be in your data. Online reviews come from the delighted and the furious; survey replies come from people with time and opinions.
- Survivorship bias. You only see what survived to be observed. Studying the habits of successful companies ignores the ones that failed doing the same things. A model evaluated only on users who stayed learns nothing about those who left.
- Non-response bias. Some chosen people never answer, and they differ from those who do. A random sample that only 10% reply to is no longer random.
The cell below builds a population of customers' monthly spend, then estimates the mean two ways at four sample sizes: honestly at random, and with a self-selection tilt where the chance of responding grows with how much you spend.
#The Bootstrap Idea
#Step 1: Resample with replacement
#Step 2: Compute the statistic
Compute whatever you care about on that one resample: the mean, the median, a correlation, a model's accuracy. This single number is one bootstrap replicate.
#Step 3: Repeat thousands of times
Do steps 1 and 2 about 1,000 to 10,000 times. You now hold thousands of replicates of the statistic. Their histogram is the bootstrap's picture of the sampling distribution.
#Step 4: Read off the answer
#Bootstrap by Hand on Ten Values
Original sample: the sum is 160, so the mean is 16.0. Sorted they read 9, 11, 12, 13, 14, 15, 16, 18, 22, 30, so the median is the average of the 5th and 6th values, (14 + 15) / 2 = 14.5. Note the 30 pulling the mean above the median.
Now draw three bootstrap resamples, each of 10 values with replacement (shown sorted, so repeats are easy to see).
| Resample | Values (sorted) | Sum | Mean | Median |
|---|---|---|---|---|
| 1 | 9, 12, 12, 12, 15, 15, 16, 16, 16, 18 | 141 | 14.1 | (15 + 15) / 2 = 15.0 |
| 2 | 9, 11, 11, 12, 14, 14, 15, 15, 22, 30 | 153 | 15.3 | (14 + 14) / 2 = 14.0 |
| 3 | 11, 14, 14, 14, 15, 16, 18, 18, 22, 30 | 172 | 17.2 | (15 + 16) / 2 = 15.5 |
Notice what happened. Resample 1 never drew the 30 or the 22, so its mean fell to 14.1. Resample 3 drew the 30 once and 22 once but missed 9, 12 and 13, and its mean rose to 17.2. Each resample is a plausible alternative dataset, and the statistic moves between them. Three replicates are far too few to read a range from, so a computer does the same thing 10,000 times. Sort the 10,000 means and cut off the lowest 250 and highest 250. The result:
- Mean: 16.0, bootstrap 95% range 12.8 to 19.9.
- Median: 14.5, bootstrap 95% range 12.0 to 19.0.
You also get a bootstrap standard error: 1.82 for the mean, close to the formula s over the square root of n, which is 6.15 divided by 3.16, or 1.94. The two agree because the formula is exactly right for the mean, up to one detail. The bootstrap treats your ten values as the whole population, and a population's standard deviation divides by n, giving 5.83 instead of 6.15. Then 5.83 divided by 3.16 is 1.84, and the bootstrap's 1.82 matches that within its own noise. The gap between 1.82 and 1.94 is the n versus n - 1 difference, not an error. For the median, the bootstrap gives you 1.75 where no tidy formula exists.
#Bootstrap in Code
choice with replacement, and two percentile calls.Look at the median histogram: it is lumpy, not smooth. With only ten values, a resample median can only be one of a handful of numbers (an observed value or the midpoint of two). That lumpiness is an honest warning that ten observations carry limited information about a median. Try replacing the ten values with 100 of your own invention and the histogram smooths out.
The demo you met earlier has a second tab for this. Scroll back to it, open the "Bootstrap" tab, resample one sample of StreamBox users 1,000 times for the mean, then switch to the median and read its interval the same way.
In a bootstrap resample of your n observations, what happens?
#Where the Bootstrap Breaks
The bootstrap is powerful and easy to over-trust. It rests on one assumption: that your sample is a decent miniature of the population. Three situations break it.
- Tiny samples. With n = 5 or 10, the sample may miss the shape of the population entirely, and resampling can only recombine what you saw. It cannot invent the rare value you did not observe. Ranges from tiny samples are too narrow and should be treated as a rough lower bound on the uncertainty.
- Heavy tails and extremes. If the population has rare, enormous values (incomes, file sizes, losses), your sample probably missed the biggest ones, so the bootstrap never shows them. It also fails outright for statistics driven by the extremes, such as the sample maximum: resampling can never produce a value larger than the largest one you already have.
- Dependent data. The resampling step treats observations as independent, shuffling them like beads in a bag. Time series, repeated measurements of the same user, and neighboring pixels are not independent. Shuffling them destroys the very structure that makes the average wobble more than the formula suggests.
The cell below shows the third failure in numbers. It builds a series where each value leans on the one before (phi = 0.8, like a slowly drifting sensor), measures how much its mean really wobbles across 2,000 fresh series, and compares with what the ordinary bootstrap claims from a single series.
#Connection to Machine Learning
#Error bars on test accuracy
A test set is a sample of the situations your model will meet. Each example is right or wrong, so score each one 1 for right and 0 for wrong. Accuracy is then the mean of these 0/1 values, and for a 0/1 column in which a fraction p are ones, the variance is p(1 - p). Check with 100 examples, 91 right: the 91 ones each sit 0.09 from the mean 0.91, the 9 zeros each sit 0.91 away, and the variance is 0.91 x 0.09 x 0.09 + 0.09 x 0.91 x 0.91 = 0.0819, which is 0.91 x 0.09.
So sigma is the square root of p(1 - p), and the standard error of an accuracy is sigma over the square root of n: SE = sqrt(p(1 - p) / n), the whole fraction under one square root. For 91% accuracy on 2,000 test examples, SE = sqrt(0.0819 / 2,000) = 0.0064, and the 95% half-width is 1.96 x 0.0064 = 0.0125, about 1.25 points (the 1.96 is the factor that makes a Normal curve 95%, close to the 2 of the 68-95-99.7 rule). On 100 examples the same 91% has SE 0.0286 and a half-width of 5.6 points. At n = 1,000 it is 1.8 points and at n = 20,000 it is 0.4 points. Same model, same true quality; only the amount of test data changed.
#Is 91.2% really better than 90.8%?
On n = 2,000 examples, a model at 90.8% has SE 0.0065 and a model at 91.2% has SE 0.0063, so each carries about 1.3 points of 95% half-width. The two scores differ by only 0.4 points, much less than that. Do not settle it by checking whether two separate ranges overlap: that is the wrong test. Whether a 0.4 point gap is real depends on the standard error of the DIFFERENCE between the two scores, and the next lesson builds exactly that test. For now, the lesson is that two bare point estimates, with no error bar at all, cannot support a ranking.
#Bagging is the bootstrap
Bagging stands for bootstrap aggregating. Train many models, each on a different bootstrap resample of the training set, then average their predictions. A single flexible model such as a deep decision tree is unstable: a small change in the data changes its answer a lot. Averaging many such models cancels that noise, exactly as averaging n observations cancels noise in a mean. A random forest is bagging plus random feature choices at each split.
#Cross-validation has variance too
A k-fold cross-validation score is an average of k scores from k different train and validation splits. Those k numbers differ for the same reason sample means differ. Report the mean and its spread across folds, not the mean alone, and remember the folds share most of their training data, so they are not independent draws and the naive formula standard deviation over square root of k understates the true uncertainty.
#Try It Yourself
Put an error bar on a 91% test accuracy three ways: the formula, the 0/1 variance, and the bootstrap, and check that the standard errors agree. The stretch step reruns the n versus n - 1 question on the ten job times.
Tests · Verify the formula table gives SE 0.0064 and a 1.25 point half-width at n = 2000 and SE 0.0286 and 5.61 points at n = 100, that the 0/1 variance equals p(1-p) = 0.0819 and s over root n matches the formula at 0.0064, that the bootstrap width is about 2.5 points at n = 2000 and about 11 points at n = 100, and that the ten job times give s = 6.15 versus 5.83 and standard errors 1.94 versus 1.84.
Running the solution prints a table where n = 100 gives SE 0.0286 and a half-width of 5.61 points, n = 1000 gives 0.0090 and 1.77 points, n = 2000 gives 0.0064 and 1.25 points, and n = 20000 gives 0.0020 and 0.40 points. The 0/1 array has variance 0.0819, exactly p(1 - p), and s over the square root of n equals the formula at 0.0064. The bootstrap on 2,000 outcomes gives a standard error of 0.0063 and a range of 0.897 to 0.922, width 2.5 points. On 100 outcomes it gives 0.0285 and a range of 0.850 to 0.960, width 11.0 points, about 4.4 times wider for 20 times less data, as the square root law says (the square root of 20 is 4.5). On the ten job times, s is 6.15 with the n - 1 divisor and 5.83 with n, so the standard errors are 1.94 and 1.84, and the bootstrap's 1.82 sits beside the second.
⚡ Playground: Probability Explorer → — explore sampling distributions and watch the Central Limit Theorem emerge.
#Key Takeaways
- A statistic is a guess, and it wobbles. The sample mean differs from sample to sample, and the sampling distribution of the mean is centered on the truth; averages are less spread out than the data they come from
- The Central Limit Theorem makes means bell-shaped. Averages of many independent draws are approximately Normal with mean mu and standard deviation sigma over the square root of n, whatever the source shape, so about 95% of sample means land within 2 standard errors of mu
- Standard error is sigma over the square root of n. In practice use s, which divides by n - 1. With sigma = 15 it goes 3.0, 1.5, 0.75 as n goes 25, 100, 400: 4 times the data for 2 times the precision, 100 times the data for 10 times
- Standard deviation is about the data; standard error is about the average. Standard deviation never shrinks with more data, standard error always does, and a report that says "plus or minus" without saying which, and n, is not usable
- Size fixes noise, not bias. A self-selected sample of 100,000 gave 84.03 with a standard error of 0.24 when the truth was 45.23; always ask how data got into the sample before trusting its error bar
- The bootstrap resamples your data with replacement to get a standard error or percentile range for almost any statistic, including the median, with no formula; it fails for tiny samples, extreme-driven statistics and dependent data
- ML uses all of it. A test accuracy of 91% on 2,000 examples is 91% give or take 1.25 points (SE = sqrt(p(1 - p) / n) = 0.0064), a 0.4 point model gap needs a proper test of the difference, bagging is the bootstrap, and cross-validation scores have a spread of their own
#Quick Check
Scores have a standard deviation of 15 points. You average n = 25 scores, then later n = 100 scores. What happens to the standard error of the average?