From Data to Distributions
After this lesson, you will be able to:
- Turn a frequency table into relative frequencies and read those as probabilities (count divided by n)
- Explain why changing the bin width of a histogram changes the story the data seems to tell
- Compute a density from a small example by hand, and say why the area of a bar is a probability while its height is not
- Build an empirical CDF by hand from 12 values and use it to answer 'what fraction is at most t?' questions
- Build a kernel density estimate by hand from three points, and explain what the bandwidth controls
- Overlay a normal curve on a histogram and say whether it looks like a good description of the data
Before You Start
#One Idea Beginners Miss
Probability textbooks start with the curve and tell you the data will look like it. Real life goes the other way: you start with data, and the curve is something you build from it. Once you see that a probability distribution is a dataset's idealised twin, the whole subject stops being mysterious.
Try it! Every code cell below runs in your browser. Edit a number, press run, and watch what changes. The best way to feel a bin width is to change it.
#From Counts to Probabilities
Run the cell. It builds the frequency table, divides by n, and then draws the same data three times with different bin widths.
The table shows 33 students in the 60s and 35 in the 70s, so a randomly chosen student has a 0.33 chance of being in the 60s and a 0.35 chance of being in the 70s. The six probabilities add to 1.00.
#Bin Width Changes the Story
- Width 2 (about 40 bars). Most bars hold just a few students, so the plot is jagged and spiky. Some of those spikes are real; most are luck. You are looking at noise.
- Width 5. A single hump with a clear centre and sensible tails. The shape is readable.
- Width 20 (a handful of bars). The detail is gone. You can see "most students are in the middle" and little else.
A teammate shows a histogram of 40 session lengths with 35 bins and says 'there are three separate groups of users, look at the three peaks'. What is the best response?
#Density Is Not Probability
Suppose you want one number for how crowded each part of the axis is, regardless of how wide you drew the bins. Counts will not do: doubling the bin width roughly doubles every count. Relative frequency has the same flaw. Work it out on a tiny example first.
Five values sit on a number line from 0 to 1: 0.05, 0.10, 0.12, 0.30 and 0.80. Cut the line into two bins of width 0.5. The bin from 0 to 0.5 holds four values and the bin from 0.5 to 1 holds one. Their relative frequencies are 4/5 = 0.8 and 1/5 = 0.2.
Redraw the same five values with four bins of width 0.25.
| bin | count | height = count / (5 x 0.25) | area = height x 0.25 |
|---|---|---|---|
| 0 to 0.25 | 3 | 2.4 | 0.6 |
| 0.25 to 0.5 | 1 | 0.8 | 0.2 |
| 0.5 to 0.75 | 0 | 0.0 | 0.0 |
| 0.75 to 1 | 1 | 0.8 | 0.2 |
That is the whole difference:
| What it is | Where the probability lives | |
|---|---|---|
| Relative frequency | a bar's share of the data | the bar's height |
| Density | share per unit of the x-axis | the bar's area (height times width) |
A probability is always an area under the density. The height is "probability per unit of x", so it only becomes a probability once you multiply by a width. When the x-axis units are small, a unit is a big slice and heights get large. Write the 100 exam scores as fractions between 0 and 1 and use bins 0.1 wide: the 70s bin then has density 0.35 / 0.1 = 3.5, while the total area is still exactly 1.
Now try the same two ideas on a bigger table. This view uses the StreamBox users, a shared synthetic table of 200 people used across Track 1, and shows their session length in minutes. Start on Histogram and drag the bin-width slider from 1 to 30 to watch the same data change its story. Then flip to Density and check that the area readout stays 1.000 at every width. Set the bin width to 10, type a = 10 and b = 30, and compare the two probabilities underneath: one counts users (115 of 200 = 0.575), the other measures area. Leave Smooth and ECDF for the next two sections, where you build them by hand.
#The Empirical CDF, Built by Hand
The histogram needs a bin width, and the answer depends on it. There is a way to describe a dataset that needs no bins at all: the empirical cumulative distribution function (ECDF). "Empirical" means built from data, and "cumulative" means a running total. It answers one question for any threshold t: "what fraction of the data is at most t?"
Do it by hand once. Here are 12 quiz scores, sorted:
| rank | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| score | 52 | 58 | 61 | 64 | 67 | 67 | 70 | 73 | 75 | 79 | 84 | 91 |
| F_hat just after | 0.083 | 0.167 | 0.250 | 0.333 | (tie) | 0.500 | 0.583 | 0.667 | 0.750 | 0.833 | 0.917 | 1.000 |
The staircase answers questions directly. F_hat(60) = 2/12 = 0.167, so about 17% of the scores are at most 60. The fraction above 75 is 1 minus F_hat(75) = 1 - 9/12 = 0.25. The fraction between 60 and 80 is F_hat(80) - F_hat(60) = 10/12 - 2/12 = 0.667.
The ECDF is the empirical version of the CDF. The theoretical CDF (covered formally in the Random Variables lesson) is a smooth S-curve; the ECDF is a staircase built from data. With more data the steps get smaller and the staircase hugs a smooth curve ever more closely. That smooth curve is the distribution the data came from.
In the 12-score example, what is F_hat(66), and what does it mean?
#Smoothing a Histogram: Kernel Density Estimates
The bell shape, phi
We write the bell shape as the function phi (the Greek letter phi, written φ). For a distance z from the middle of the bell:
φ(z) = exp(-z²/2) / √(2π)
You do not need to compute that by hand. Read it off a few values: φ(0) = 0.399 is the top of the bell, φ(1) = 0.242 one step away, φ(2) = 0.054 two steps away and φ(3) = 0.004 three steps away. It is symmetric (φ(-1) is also 0.242), and the area under the whole bell is exactly 1.
Three points, worked by hand
Here is each bell's height at the spots t = 1 to 6, then the sum, then the sum divided by 3:
| t | bell at 1 | bell at 3 | bell at 4 | sum | sum / 3 |
|---|---|---|---|---|---|
| 1 | 0.399 | 0.054 | 0.004 | 0.457 | 0.152 |
| 2 | 0.242 | 0.242 | 0.054 | 0.538 | 0.179 |
| 3 | 0.054 | 0.399 | 0.242 | 0.695 | 0.232 |
| 4 | 0.004 | 0.242 | 0.399 | 0.645 | 0.215 |
| 5 | 0.000 | 0.054 | 0.242 | 0.296 | 0.099 |
| 6 | 0.000 | 0.004 | 0.054 | 0.058 | 0.019 |
Read the row t = 3. The bell centred at 3 is at its top, 0.399. The bell at 4 is one step away, so it contributes φ(1) = 0.242. The bell at 1 is two steps away, so it contributes φ(2) = 0.054. The three heights add to 0.695.
Why divide by 3? Each bell has area 1, so three bells have area 3. Dividing by n = 3 brings the total area back to 1 (a numerical check gives 1.000). The last column is the KDE.
Look at its shape. It is tallest near t = 3 and 4, where two bells overlap, and lower near t = 1, where only one point sits. Where points crowd together the curve rises, and where they are sparse it falls.
What does the bandwidth do? With h = 2 each bell would be twice as wide and half as tall, so its area stays 1. That "half as tall" is the 1/h in the formula below. A small h hugs every point, and a large h smears everything into one hill. It is the same trade-off as the bin width.
Now the formula. It says exactly what the table did: for each of the n data points x_i, measure the distance from t in steps of h, read the bell height, add them all up, and divide by n and h.
gaussian_kde produces exactly the same curve as the hand-built version.A KDE is the first place where the "data becomes a distribution" idea turns concrete. You handed it 100 numbers and got back a smooth density you can integrate and compare with theory. It still makes no assumption that the data is normal, or anything else. It is a flexible picture of where the data lives.
#Where This Leads: Fitting Named Distributions
Here is a quick preview with no new maths. Draw a normal curve with the same centre (mean) and spread (standard deviation) as the 100 scores on top of their histogram, and ask a plain question: does the curve look like the bars?
The curve follows the bars closely. The share of scores from 60 up to 80 is 0.68 in the data and about 0.68 under the curve, so a bell with two numbers, 69.65 and 10.00, describes these 100 scores well. Not every column is that kind. The session lengths in the StreamBox view are lopsided, and a bell would fit them less well.
Two later lessons pick this up. The Distribution Zoo is a menu of named shapes and how to choose among them. Sampling, Standard Error & the Bootstrap asks how much a picture like ours would change if you collected another batch of data. In machine learning, a training set plays the part of the 100 scores: a pile of examples, and a model is something fitted to the pattern in them.
#Try It Yourself
You will rebuild the lesson's four pictures as numbers on the same 100 exam scores: relative frequencies, a density, the ECDF, and a KDE value. Each step has a TODO. Write your answer under it, run the cell, and compare with the solution.
Tests · Verify the six relative frequencies are 0.02, 0.14, 0.33, 0.35, 0.14, 0.02 and sum to 1; the tallest density height is 0.035 while the total area is 1; ecdf(75) is 0.73, the share above 85 is 0.04, and the share from 60 up to 80 is 0.68; the KDE values at t = 70 are about 0.0391, 0.0377 and 0.0222 for h = 1, 4 and 15.
The real values: the counts are 2, 14, 33, 35, 14 and 2, the tallest density height is 0.035 while the total area is 1.0, the ECDF at 75 is 0.73, and the KDE at 70 is 0.0391, 0.0377 and 0.0222 for h = 1, 4 and 15. The three KDE values shrink as h grows because a wider bell spreads the same area over more of the axis, so the peak gets lower.
#Key Takeaways
- A relative frequency (count divided by n) is an empirical probability, and a table of them is the empirical distribution of the data
- Bin width is a choice, not a fact: too narrow shows noise as structure, too wide hides real structure, so check that a feature survives changes in bin width before believing it
- Probability is an area under a density, never the height: a narrow bar can be taller than 1, but its area is at most 1 and all the areas add up to exactly 1
- The empirical CDF needs no bins: F_hat(t) is the fraction of the data at most t, a staircase that climbs from 0 to 1
- A kernel density estimate puts one smooth bell on every data point, adds them and divides by n, with the bandwidth playing the role of bin width
- Fitting a named distribution to these pictures is the next step, and the Distribution Zoo and Sampling lessons pick it up
#Quick Check
In a dataset of 200 values, 50 fall in the bin from 20 to 30. What is the relative frequency of that bin, and what does it estimate?