The last lesson gave you mean, median, standard deviation and the 68-95-99.7 rule. Those are perfect for a tidy bell curve. Real data is rarely tidy: delivery times have a long late tail, salaries have a billionaire, latencies have a 99th-percentile spike. This lesson gives you the tools for that world: quantiles, box plots, skewness, and robust statistics that refuse to be bullied by one bad row.
Learning Objectives
After this lesson, you will be able to:
Compute any quantile (percentile, quartile) by hand using the same interpolation rule NumPy and Excel use, and read a box plot step by step
Use the IQR and the 1.5 x IQR fence to flag outliers, and explain why the z-score rule can miss them
Compute a z-score, then use z-scores to compute skewness by hand and read what its sign and size mean
Decide what to do with an outlier (error, rare-but-real, or the signal) and choose robust tools: median, MAD, trimmed mean, log transform
Explain why ML cares: p50/p95/p99 latency, standard vs robust scaling, clipping, and MSE vs MAE vs Huber loss
Every example here uses the same synthetic dataset: 40 food-delivery times in minutes. Most trips take 20-37 minutes, but four were badly late (52, 61, 74 and 95 minutes), and one was suspiciously fast (17). The data was generated with a seeded random generator and is printed in full below so you can check every number in this lesson yourself.
Run the cell below first. It sorts the data and draws a histogram with the mean and median marked. Notice the long right tail and the gap between the two lines. Everything that follows is about how to describe, measure and handle exactly this kind of data.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The mean is 32.08 minutes and the median is 29. The standard deviation is 14.52, which is huge compared with how tightly most trips cluster. Keep that gap in mind; the rest of this lesson gives you the tools to describe, measure and handle exactly this kind of data.
If you have used Excel's PERCENTILE.INC function, you have already used the rule NumPy follows by default. See it on five sorted values: 10, 20, 30, 40, 100.
Five values leave four gaps between them. Walk from the smallest value (0% of the way) to the largest (100%), and each gap is worth 25% of the journey.
Value
10
20
30
40
100
Fraction of the way along
0%
25%
50%
75%
100%
So the 25th percentile is exactly the second value, 20. The median is 30 and the 75th percentile is 40. Nothing to blend. Now ask for the 30th percentile. It sits between the 25% value (20) and the 50% value (30), and 30% is 5 points into that 25-point gap, which is one fifth of the way. So p30 = 20 + 0.2 x (30 - 20) = 22. Excel's PERCENTILE.INC and np.percentile([10, 20, 30, 40, 100], 30) both return 22.
That is the whole idea: find which two neighbours the fraction lands between, then move that fraction of the way from one to the other.
With 11 values there are 10 gaps, so each gap is worth 10% of the way, and the 25th percentile falls 2.5 gaps in: halfway between two values. Written as steps:
Sort the data and number the positions from 0. Computers count from 0, so the smallest value is position 0 and the largest is position n - 1 (here 10).
Multiply: h = (n - 1) x q. This is "how many gaps along", where q is the fraction you want (0.25 for the 25th percentile).
Split h in two. The whole-number part (h rounded down, written floor(h)) picks the lower neighbour. The decimal part says how far to move toward the next value up.
Blend: result = lower neighbour + decimal part x (next value - lower neighbour).
Worked example with 11 values. Take the first eleven distinct trips, sorted: 14, 17, 19, 22, 24, 26, 28, 31, 35, 41, 78. The strip below lays them out by position and marks where each quantile lands.
position 0 1 2 3 4 5 6 7 8 9 10
value 14 17 19 22 24 26 28 31 35 41 78
Q1 = 20.5 ^ ^ halfway between 19 and 22
median = 26 ^
Q3 = 33 ^ ^ halfway between 31 and 35
p90 = 41 ^
Positions run from 0 to 10, so n - 1 = 10.
Quantile
h = 10 x q
Lower neighbour
Next up
Result
Q1 (q = 0.25)
2.5
x[2] = 19
x[3] = 22
19 + 0.5 x (22 - 19) = 20.5
Median (q = 0.50)
5.0
x[5] = 26
none needed
26
Q3 (q = 0.75)
7.5
x[7] = 31
x[8] = 35
31 + 0.5 x (35 - 31) = 33
p90 (q = 0.90)
9.0
x[9] = 41
none needed
41
Notice Q1 = 20.5 is not one of the original numbers. That is normal: interpolation produces a value between data points. The "fraction of the way" is just the decimal part of h.
That is the whole recipe. Here it is as one formula, with the symbols in the same order as the four steps:
The interquartile range is Q3 minus Q1. It is the width of the middle 50% of the data, and it ignores the extreme 25% at each end, so it cannot be dragged around by outliers. For the 11-value example, IQR = 33 - 20.5 = 12.5. For all 40 delivery times, Q1 = 25.75, Q3 = 31, so IQR = 5.25 minutes, compared with a standard deviation of 14.52. The IQR describes the crowd; the standard deviation got inflated by the late tail.
The whiskers stretch to the most extreme data points still inside the fences. Here the lower whisker reaches 20 and the upper whisker reaches 37.
Every dot beyond the fences is flagged individually. Here: 17 (below 17.875), and 52, 61, 74, 95 (above 38.875).
Look at the asymmetry. All but one of the flagged dots are on the upper side, and the median sits nearer the bottom of the box than the top. Lopsidedness like this is what skew looks like.
lower fence=Q1−1.5IQR,upper fence=Q3+1.5IQR
Now see a box plot move. This one runs on the 200 StreamBox users, a shared synthetic table reused across Track 1 (the 40 delivery times in this lesson stay as they are). Drag the extreme-user slider up and watch the mean and standard deviation jump while the median barely moves; the table also shows the MAD, a robust spread measure explained later in this lesson. Then change k, the 1.5 in the fence rule, to see which dots the fence flags.
Loading visualization...
What Do You Think?
In the 40 deliveries, the mean is 32.08 and the median is 29. Which statement best explains the gap?
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
Quick check
Q1 = 40 and Q3 = 60 for some dataset. What are the upper and lower outlier fences?
The Descriptive Statistics lesson showed that the median resists outliers. Here is that idea on the deliveries, where the mean is 32.08 and the median is 29. Suppose the 95-minute trip had been logged as 950 by a typo. The median stays exactly 29, because it only uses the ordering of the values and 950 is still the biggest. The mean leaps from 32.08 to 53.45, and the standard deviation from 14.52 to 145.76. Quartiles and the IQR do not move either.
So a mean above the median is a fingerprint of shape:
Mean and standard deviation can be identical for two datasets that look nothing alike. Skewness is one extra number that says which way the tail leans and how strongly. It is built from z-scores, so we start there.
A z-score says how many standard deviations a value sits above or below the mean: z = (value - mean) / standard deviation. A z-score of +2 means "two standard deviations above the mean", -1 means "one below", 0 means "exactly at the mean". Because the units cancel, z-scores let you compare things measured in different units.
Take five values: 2, 3, 4, 5, 16.
The mean is 30 / 5 = 6.
The deviations from the mean (value minus mean) are -4, -3, -2, -1 and +10.
Square them: 16, 9, 4, 1, 100. They add to 130. Divide by 5 to get the variance, 26. Its square root is the standard deviation, written with the Greek letter sigma (σ): σ = 5.10. (Skewness divides by n, so this is the divide-by-n version of the standard deviation.)
Cube each z-score (multiply it by itself three times).
Average the cubes. Add them up and divide by n.
Read the sign. Positive means a long right tail, negative a long left tail, near zero means roughly symmetric.
Read the size. The bigger the number, the more lopsided.
Value
z-score
z squared
z cubed
2
-0.78
0.62
-0.48
3
-0.59
0.35
-0.20
4
-0.39
0.15
-0.06
5
-0.20
0.04
-0.01
16
+1.96
3.85
+7.54
Why cube? Cubing keeps the sign. A negative number multiplied by itself three times stays negative: the first two factors make a positive, and the third negative factor flips it back. For z = -0.78 the cube is about -0.48. So the cubes of values below the mean are negative and the cubes of values above are positive. Squaring would not do: the "z squared" column is all positive, so it cannot tell a left tail from a right one. Cubing also makes far-away values count for much more: the 16 contributes +7.54 while the four small values together contribute only -0.75.
The cubes add to 6.79, and the average is 6.79 / 5 = 1.36. The sign is positive and the size is above 1: a clear long right tail, which matches the single 16 stretched out to the right. Here is the same recipe as one formula.
skew=n1i=1∑n(σxi−xˉ)3
For the 40 deliveries, skewness is 2.91: very strongly right-skewed, exactly what the histogram showed. Skewness tells you which side the tail is on, not how many outliers there are, and on small samples it is noisy, so always look at the histogram too.
1. The IQR fence (above) flagged five of our 40 trips: 17, 52, 61, 74, 95. It is built from quartiles, so the outliers themselves cannot distort the fence.
2. The z-score rule flags values with |z| above 3, using the mean and standard deviation. For the delivery data the mean is 32.08 and the standard deviation is 14.52, so the cutoff is 32.08 + 3 x 14.52 = 75.6. Only 95 (z = 4.33) is flagged. The 74-minute trip has z = 2.89 and slips under the line, as do 52 and 61.
This is called masking: the outliers inflate the standard deviation, which widens the z-score band, which hides the outliers. The IQR fence does not suffer from this. This is why robust methods matter: the ruler you measure outliers with must not be bent by the outliers.
A flag is a question, not a verdict. Ask what produced the point:
The outlier is...
Example in our data
What to do
An error (broken sensor, typo, unit mix-up)
A "-5 minutes" delivery, or 3,000 because someone typed seconds
Fix it or drop it, and log that you did
Rare but real (genuine, just unusual)
The 95-minute trip during a storm
Keep it. It is part of reality your model must handle; use robust methods rather than deleting it
The signal (the thing you are actually looking for)
A card transaction 40x the usual amount; a server latency spike
Do not remove it. Fraud and fault detection are outlier detection
In our data, ask: is the 17-minute trip an error (impossibly fast) or a one-block delivery? The fence flags it, but only domain knowledge tells you. Deleting every flagged point is a quiet way of making a model look better by pretending the hard cases do not exist.
#Robust Statistics: Summaries That Refuse to Be Bullied
A statistic is robust if a few extreme values cannot move it much. Three to know:
A trimmed mean sorts the data, throws away a fixed fraction from each end, and averages the rest. Trimming 10% from each end of the 40 deliveries (4 values each side) gives 28.78 minutes, compared with the plain mean of 32.08 and median of 29. It keeps most of the mean's efficiency while ignoring the tails.
Recall from the Math Toolkit lesson that a log shrinks big numbers more than small ones. The natural log turns 10, 100 and 1,000 into 2.30, 4.61 and 6.91: every time the number is multiplied by 10, the log goes up by the same step (2.30). That is exactly the medicine a long right tail needs.
Many real quantities are produced by multiplying effects together (prices, durations, traffic, populations), not adding them. Their distribution has a long right tail. Taking logarithms turns multiplication into addition, which tends to produce a more symmetric, bell-like shape.
For the delivery data, apply the natural log to every trip time (np.log(times)). Skewness falls from 2.91 to 1.78. It is not perfect because our late outliers are extreme, but the transform pulls the tail in. The log also changes how distances work: on the log scale, 30 to 60 minutes (doubling) is the same step as 60 to 120. That is usually the right way to think about "twice as late".
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
Quick check
You compute the z-score rule (|z| above 3) on a dataset with a few huge outliers, and it flags almost nothing. What is the most likely explanation?
Service speed is reported as percentiles, never the mean. The name p95 means the 95th percentile: the value that 95% of requests beat. For our 40 deliveries:
p50 = 29 min (the typical trip)
p95 = 61.65 min (95% of trips were faster)
p99 = 86.81 min (99% of trips were faster)
If the app promised "30 minutes" based on the median, 14 of the 40 trips (35%) would have broken it. A promise a company can keep, such as "within 62 minutes for 95% of orders", is a quantile statement. Teams monitor p99 because it is where the pain concentrates: a request that is slow at p99 for one call becomes the typical experience when a page makes 100 calls. For LLM serving, time-to-first-token p95 and p99 are the headline reliability numbers for the same reason.
Scaling means rewriting each feature on a common footing before a model sees it. The last lesson introduced standardization, z = (x - mean) / std, which is the z-score you met above. The robust sibling swaps in the median and IQR (scikit-learn calls it RobustScaler):
zstandard=sx−xˉ,zrobust=IQRx−median(x)
On the delivery data, standard scaling gives the 95-minute trip z = 4.33 and squashes the crowd: the ordinary trips between 24 and 35 minutes land in a band only 0.76 wide. Robust scaling gives the 30-minute trip 0.19 and the 95-minute trip 12.57, and the same crowd spans 2.10, so it keeps a meaningful spread while the outlier is plainly far away. If a model must treat outliers as outliers, use the robust scaler. If your data is clean and roughly bell-shaped, standard scaling is fine and slightly more efficient.
Clipping, which statisticians call winsorizing, caps values at chosen quantiles before training. Clip delivery times at the 5th and 95th percentile (20.95 and 61.65) and the mean drops from 32.08 to 31.05, without deleting any rows. It is the middle path between "keep a destructive outlier" and "delete it". Gradient norms get the same treatment in deep-learning training: gradient clipping caps the size of an update so one bad batch cannot throw the weights off a cliff.
A loss function scores how wrong a model's prediction is, and training tries to make the total score small. The choice of loss is a choice about outliers. Suppose a model predicts a constant delivery time c, and the error on one trip is e = actual - c.
MSE (mean squared error) squares each miss, then averages. A 65-minute miss costs 65^2 = 4,225, over four thousand times more than a 1-minute miss. The best constant under MSE is the mean.
MAE (mean absolute error) averages the size of each miss. Costs grow in a straight line. The best constant under MAE is the median.
Huber loss (named after the statistician Peter Huber) is quadratic for small misses (below a threshold called delta) and linear for big ones. It behaves like MSE near the optimum, smooth and easy to optimise, and like MAE for outliers.
Lδ(e)={21e2δ(∣e∣−21δ)∣e∣≤δ∣e∣>δ
Before and after, on five values. Predict the constant 30 for five trips. First the trips are 28, 29, 30, 31, 32 (no outlier). Then the last one becomes 95, giving 28, 29, 30, 31, 95. The errors in the second case are -2, -1, 0, +1, +65. Squared: 4, 1, 0, 1, 4,225, which average to 846.2. Absolute: 2, 1, 0, 1, 65, average 13.8. Huber with delta = 5: 2, 0.5, 0, 0.5, 312.5, average 63.1.
Loss (constant = 30)
No outlier
With one outlier
Grew by
MSE
2.0
846.2
423 times
MAE
1.2
13.8
11.5 times
Huber (delta = 5)
1.0
63.1
63 times
One bad trip makes the MSE 423 times bigger, but the MAE only 11.5 times bigger. Huber sits between. That is why a model trained on MSE is pulled hard toward the outlier and a model trained on MAE or Huber is not. (Huber here uses 0.5 x e^2, so its no-outlier value is half of MSE.)
This is the theme of the lesson, applied: mean/MSE and median/MAE are two answers to "how much should an outlier count?" Choose deliberately. If rare extremes are your business (a big forecast miss is genuinely terrible), MSE is right. If they are noise, MAE or Huber is.
Tests · Verify Q1=25.75, median=29, Q3=31, IQR=5.25, flagged values [17, 52, 61, 74, 95], p95 about 61.65, p99 about 86.81, skew about 2.91, and that the answer names the median.
Run the solution and you should see Q1=25.75, median=29.0, Q3=31.0, IQR=5.25 on the first line, fences=[17.875, 38.875] with flagged=[17 52 61 74 95] on the second, p95=61.65 and p99=86.81 on the third, skew=2.91, and a closing sentence naming the median.
Quantiles come from sorting and interpolating. The position is h = (n - 1) x q; blend the two neighbouring sorted values by the fractional part. Quartiles, percentiles and the median are all quantiles
The IQR (Q3 - Q1) measures the spread of the middle 50% and ignores the extremes. The 1.5 x IQR fence and the box plot are built on it, and unlike the z-score rule it cannot be masked by the outliers it is hunting
Under right skew the mean sits above the median. Delivery times, salaries and latencies are all like this; report the median and percentiles, not the mean
Skewness is the average of the cubed z-scores, and it says which way the tail leans. Zero means symmetric, positive means a long right tail, and it is noisy on small samples. Kurtosis, an optional extra, measures tail heaviness the same way
Detecting an outlier is arithmetic; deciding about it is judgment. Ask whether it is an error, rare but real, or the signal, and never delete a point just because it is inconvenient
Robust tools exist for a reason. The median, MAD, trimmed mean and log transform resist extremes, and in ML they appear as p95/p99 monitoring, robust scaling, clipping, and MAE/Huber loss
Sorted data: 14, 17, 19, 22, 24, 26, 28, 31, 35, 41, 78 (n = 11). Using NumPy's linear interpolation, what is Q1 (the 25th percentile)?
Next up: From Data to Distributions. You now know how to describe one real dataset honestly. Next you learn how to go from the data you see to the distribution that probably generated it.
Your Reflection
Saves automatically
What’s one thing you learned? What’s still confusing?