Before any model fits anything, someone calls df.describe(). Mean, median, std — a handful of summary numbers reveal most common data problems: skew, outliers, scale mismatches, broken pipelines. The fanciest deep network can't save you from skipping this step. Statistics isn't homework — it's the first line of defense against silently broken ML.
Learning Objectives
After this lesson, you will be able to:
Calculate the average, middle value, most common value, and spread of a dataset, and know when to use each one
Describe the bell curve by its mean and standard deviation, work out a variance by hand, and convert a score into a z-score
Understand why the bell curve shows up everywhere in nature and AI, thanks to a powerful math rule called the Central Limit Theorem
Apply the 68-95-99.7 rule to identify outliers and perform z-score standardization, the preprocessing step that makes gradient descent converge faster
Statistics is not about math — it is about storytelling with numbers. Every time you check your average screen time, compare your test scores to the class, or wonder if a game's matchmaking is fair, you are already doing statistics. This lesson just makes you precise about it.
This lesson covers the foundational statistics you need before diving into probability, inference, and machine learning. If you can add, subtract, and find a square root, you are ready. Every new word and symbol is explained the first time it appears, and every formula comes after a small example worked with real numbers.
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
Start with three test scores: 70, 80 and 90. Add them up: 70 + 80 + 90 = 240. Divide by how many scores there are: 240 / 3 = 80. That number, 80, is the mean (the everyday "average").
It is the "balance point" of the data: put the three scores on a seesaw and it balances exactly at 80, because 70 is 10 below it and 90 is 10 above it.
Now the formula, which says the same thing in symbols. Write n for how many values you have, and x₁, x₂, … for the values themselves. The mean is written x-bar (a letter x with a bar over it): x-bar is the average of the sample you actually have. The big symbol Σ (capital sigma) simply means "add these up". It is a different symbol from the lowercase σ you will meet shortly.
The median is the middle value when you sort the data. For 70, 80, 90 the middle one is 80. Half the values are below it, half are above. Unlike the mean, the median is not affected by extreme values.
What Do You Think?
If the mean salary at a company is $500K, does that mean most employees earn around $500K?
Not necessarily. If the CEO earns $10M and 19 employees each earn $50K, the mean salary is ($10M + 19 x $50K) / 20 = $547,500. But the median is $50K, and that represents the typical employee far better. A few huge values drag the mean toward them; the median does not budge. When the values trail off into a long tail on one side like this, the data is called skewed, and the median is the honest summary. That is why economists report median household income and real estate listings show the median home price. If the mean and median are far apart, treat that gap as a skew alarm.
The mode is the value that appears most frequently. It is especially useful for categorical data (like "what genre of music do students prefer?") where computing a mean makes no sense.
Quick check
Which summary statistic should you report for the typical home price in a US city?
#Measures of Spread: How Spread Out Are the Scores?
Knowing the center is only half the story. Two datasets can have the same mean but look completely different. Imagine two classes both averaging 75: in one class, everyone scored between 70-80; in the other, scores ranged from 20 to 100. You need a measure of spread.
We want one number that says "how far, typically, are the values from the mean?" Work it out by hand for the same three scores, 70, 80 and 90, whose mean is 80.
Score
Distance from the mean (80)
Distance squared
70
70 − 80 = −10
100
80
80 − 80 = 0
0
90
90 − 80 = 10
100
Total
0
200
Look at the middle column: the distances add up to exactly 0, because the mean is the balance point and the pulls on each side cancel. Adding raw distances would always give 0 and tell us nothing. Squaring each distance fixes that: every square is positive, and big distances count for more. The squares add up to 200.
Now average the squares, which means dividing the 200:
Divide by 3 (the number of scores) and you get 66.7. Use this when your three scores are the whole group you care about, such as every student in the class.
Divide by 2 (one fewer than the number of scores) and you get 100. Use this when your three scores are only a sample, a part you collected from a bigger group. Almost all ML data is a sample, so this is the one you will use most often. The Bessel's-correction DeepDive below explains the "one fewer".
Either result is the variance: the average squared distance from the mean. Its units are squared (points squared, not points), which is awkward. So take the square root to get back to the original units: the square root of 66.7 is 8.2, and the square root of 100 is 10. That is the standard deviation, the typical distance from the mean.
Now the symbols. Here are four of them in words before you see the formula:
μ (mu) is the Greek letter for the true average of the whole group.
x-bar is the average of the sample you actually have. It is our estimate of μ.
σ (sigma, lowercase) is the true standard deviation of the whole group; σ² is the true variance.
s is the standard deviation you compute from your sample, and s² is the sample variance.
σ2=n1i=1∑n(xi−μ)2s2=n−11i=1∑n(xi−xˉ)2
Read the left formula as "the average of the squared distances from μ". The right formula is the same recipe with x-bar and with n−1 in place of n. In code, ddof stands for "delta degrees of freedom", which is just the number you subtract from n before dividing: ddof=1 means divide by n−1.
A standard deviation gives you a ruler for "how unusual is this value?". Suppose a class has a mean of 80 and a standard deviation of 5, and you scored 90.
You are 90 − 80 = 10 points above the mean.
One standard deviation is 5 points, so 10 points is 10 / 5 = 2 standard deviations above the mean.
That count, 2, is your z-score. A z of 2 means "two standard deviations above average". A z of 0 means exactly average, and a z of −1 means one standard deviation below. The z-score lets you compare things measured in different units: 2 standard deviations above on a maths test and 2 above on a typing test are equally unusual.
For data that follows a bell-shaped (normal) distribution, z-scores connect to probabilities through a powerful shortcut:
About 68% of values have a z-score between −1 and 1, which means they fall within 1 standard deviation of the mean
About 95% fall within 2 standard deviations
About 99.7% fall within 3 standard deviations
This means if the average test score is 75 with a standard deviation of 10, about 95% of students scored between 55 and 95. For normal data, anything outside 3 standard deviations is rare (about 0.3%), but the rule applies only to normal distributions. Back in our class (mean 80, standard deviation 5), a score of 90 has z = 2, so it sits right at the edge of the middle 95%.
df.describe() is the first thing a working ML engineer types on a new dataset. The cell below builds the core of that output by hand from a class of test scores: the mean, median, mode, standard deviation and range. It then plots the histogram and overlays the mean, the median and a one-standard-deviation band so you can see whether mean and median diverge. Note ddof=1: it is the "divide by n−1" sample version you just learned.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
Run as-is and you should see n = 30, mean = 81.133, median = 81.500, mode = 88 (appears 2 times), std = 8.982 and a range of 33 (63 to 96). Mean and median are close, so the scores are roughly symmetric. Now append 5, 8 and 12 and re-run: the mean falls to 74.515 while the median only slips to 80.0.
#Probability Distributions: The Shape of Randomness
A probability distribution describes how likely each possible outcome is. It is the "shape" of your data. Different real-world processes produce different shapes, and recognizing the shape tells you a lot about the underlying process. One shape matters more than all the others.
The normal distribution is the most famous distribution in all of statistics. You have seen it in class: the symmetric bell curve where most values cluster near the mean and extreme values are rare on both sides. It is completely described by two numbers you already know: the mean μ (where the center sits) and the standard deviation σ (how wide the bell is). A bigger σ gives a wider, flatter bell; a smaller σ gives a taller, narrower one. The 68-95-99.7 rule above is a statement about exactly this shape.
f(x)=σ2π1e−2σ2(x−μ)2
Nobody evaluates this formula by hand. It is here so you can see that μ and σ are the only two dials. Why does the normal distribution show up everywhere? Heights, test scores, measurement errors, stock returns (roughly) — the answer is the Central Limit Theorem, which we will get to shortly.
The viz below shows Normal, Binomial and Poisson side by side. Pick a tab, move the sliders, and read the Mean, Variance and Std Dev tiles under the chart.
Loading visualization...
Try this: On the Normal tab, set μ = 0 and σ = 1 (the "standard normal"), then raise σ and watch the bell flatten while the Std Dev tile follows your slider. Set the shading sliders to a = −1 and b = 1 and read the probability: about 0.68, the "68" in the rule. Widen to −2 and 2 for about 0.95. Optional, once you have read the DeepDive above: switch to the Binomial tab, set n = 20 and p = 0.5 for a symmetric, bell-like bar chart, then slide p down to 0.05 to see an extreme right skew, because almost every trial fails. On the Poisson tab, set λ = 1 for another right-skewed shape. The shape of a distribution is what the mean and standard deviation cannot tell you on their own; the next lesson gives you numbers for it.
This second explorer offers four other shapes: Normal, Uniform, Exponential and Beta. Pick one with the Distribution menu; each has its own sliders.
Try it: Adjust the parameters and watch the distribution changeInteractive
Loading visualization...
Try this: Start with Normal, μ = 0 and σ = 1. Raise σ and watch the bell flatten. Switch on Show Samples to see random draws pile up under the curve, then drag Range Start and Range End to shade a region under the curve. Now choose Exponential: the curve has a long right tail, a skewed shape like the CEO-salary data. Compare its Mean with the point where the curve peaks (the mode) and notice they are far apart, the same warning sign you saw for the mean and median.
Take any distribution — flat, lopsided, even a completely weird custom one. It does not matter what shape it is. The Central Limit Theorem works on all of them.
Draw a random sample of n values from that distribution. Compute the mean of your sample. This is one "sample mean." It is a single number — the average of your random draw.
Do this again and again. Each time, draw n fresh values and compute the sample mean. After hundreds or thousands of repetitions, you have a collection of sample means.
Plot the distribution of those sample means. No matter what the original distribution looked like, the distribution of sample means will be approximately normal (bell-shaped). The larger n is, the more perfectly normal it becomes. This is the Central Limit Theorem — the single most important result in statistics.
In plain words: average enough independent random values and the averages form a bell curve, whatever shape the individual values had. That is why the normal distribution shows up everywhere: a person's height, a test score, a measurement error are each the combined effect of many small independent contributions. Real data is not always normal; it is averages and sums that tend toward normal. The Sampling, Standard Error & the Bootstrap lesson later explains why this happens and how wide the bell is.
See it happen. Pick a skewed source in the viz below, draw batches of 30 values, and watch the batch averages pile into a bell shape even though the source is lopsided.
Before training any model, you compute descriptive statistics on every feature: mean, median, standard deviation, min and max. These numbers tell you whether features are on the same scale (if not, you need normalization), whether there are outliers (which might need clipping), and whether the mean and median agree (a first hint that a feature is roughly symmetric, which some algorithms assume).
Many ML algorithms (gradient descent, KNN, SVMs) are sensitive to the scale of features. Standardization subtracts the mean and divides by the standard deviation: z = (x - mean) / std. This is the z-score you computed by hand earlier, applied to every value in a feature, and it gives every feature mean 0 and standard deviation 1. You cannot do this without knowing the mean and standard deviation first.
Values more than 3 standard deviations from the mean are rare (the 68-95-99.7 rule). This is the simplest anomaly detector: compute the mean and standard deviation, flag anything beyond 3 sigmas. Credit card fraud detection, server monitoring, and quality control all start here.
Linear regression assumes normally distributed errors. Naive Bayes often assumes Gaussian features. Understanding what distribution your data follows, and what happens when it does not — is the difference between a model that works and one that silently fails.
Now compute everything from this lesson on the 30 class scores, with no library hiding the arithmetic. Fill in the TODOs, run, and compare with the solution. Note that the standard deviation must divide by n−1, because these 30 scores are a sample.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
Tests · Verify the printed values: Mean 81.13, Median 81.50, Std Dev (n-1) 8.98, matches np.std(ddof=1) True, Std Dev with n 8.83, within 1 std 19 of 30 (63.3%), within 2 std 29 of 30 (96.7%), z of 96 is 1.66, z of 63 is -2.02, and with a 15 the mean is 79.00 (moved -2.13) while the median is 81.00 (moved -0.50).
Check your output against these real values. The mean is 81.13 and the median 81.50, close together, so the scores are roughly symmetric. The standard deviation is 8.98 with n−1; dividing by n would give 8.83, the smaller number you get from np.std(scores) with its default. Of the 30 scores, 19 (63.3%) sit within 1 standard deviation and 29 (96.7%) within 2: a little under 68% and a little over 95%, which is about what the rule predicts for 30 values that are only roughly bell-shaped. The top score, 96, has z = 1.66 and the bottom score, 63, has z = −2.02, so the lowest score just sits outside the 2-standard-deviation band. Finally, one score of 15 drags the mean down by 2.13 (to 79.00) but moves the median by only 0.50 (to 81.00).
Mean, median, and mode measure the center differently. The mean is sensitive to outliers, the median is robust, and the mode identifies the most common value; always check whether mean and median diverge, which signals skewed data
Variance is the average squared distance from the mean (divide by n−1 for a sample), and standard deviation is its square root. It measures consistency: small means tightly clustered data, large means widely spread
A z-score counts how many standard deviations a value is above or below the mean: z = (x − mean) / std
For bell-shaped data the 68-95-99.7 rule gives instant intuition: about 68%, 95% and 99.7% of values fall within 1, 2 and 3 standard deviations of the mean. It applies only to normal distributions
The normal distribution is defined by its mean and standard deviation, and the Central Limit Theorem explains why it is everywhere: averages of many independent values form a bell curve whatever shape the individual values had
Three scores are 2, 4 and 6 (mean 4). Treating them as a sample, what is the standard deviation?
Next up: Quantiles, Shape & Reading Real Data. You will learn how to cut a dataset into slices, read a box plot, measure skew, and decide what to do with outliers.