What’s one thing you learned? What’s still confusing?
Causation, Confounding & Simpson's Paradox
A strong correlation can appear with no causal link at all. This lesson shows how a confounder manufactures one, what randomisation removes, how Simpson's paradox reverses a pooled table inside every subgroup, and how selection bias and shortcuts make models learn the wrong thing.
The Distribution Zoo, Part 1: The Core Distributions
Every dataset has a shape, and every loss function hides a distribution. Meet the handful of distributions that cover most of ML (Uniform, Exponential, Beta, Dirichlet, Categorical, Multinomial and Student-t), each with its story, its numbers and its use.
The Distribution Zoo, Part 2: Fitting Data and Heavy Tails
Real data is positive, skewed and has a tail. This lesson adds the Lognormal, shows how to check any distribution with log-likelihood, AIC, Q-Q plots and tests, covers Gamma, Chi-square and power laws, and ends with a decision guide and a worked fit of 60 API response times.
Interactive Labs for This Track
Loss Landscape
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Vectors & Matrix Operations
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Probability Distributions
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Ask questions, share insights
Here is where this two-lesson story ends up. Three real-looking results that mislead, each of which you will be able to catch by the end of Part 2:
Correlation answers one narrow question: if I know X, how well can I guess Y with a straight line? It says nothing about why, and nothing about what happens if you change X.
Cov(X, Y) = E[(X − μ_X)(Y − μ_Y)], where μ (mu) is the mean of a variable. This lesson takes the standardised version, Pearson's r, and puts it under stress. Causal estimation itself lives in the Causal Inference Basics lesson in the classical ML track, which covers propensity scores (the modelled chance that a unit receives the treatment, used to compare like with like) and Double ML (using machine-learning models to strip the confounders out of both variables before comparing them). Here we stay with the foundations: what the number is, how it breaks, and how to avoid being fooled by it. This is Part 1 of two lessons, and it follows Checkpoint: Foundations Review. Part 2, Causation, Confounding & Simpson's Paradox, takes the number into the question of cause.Every section below uses numbers you can verify by hand or by running the code. All counts and simulations are synthetic and labelled as such. The one real dataset is Anscombe's quartet, a famous public teaching example from 1973.
Covariance measures whether X and Y sit above or below their means together:
The trouble with covariance is its units. Heights in centimetres and weights in kilograms give a covariance in cm·kg, and changing centimetres to metres changes the number by a factor of 100 without changing the relationship. Dividing by both standard deviations removes the units:
How to read the value:
r(X, Y) = r(Y, X)) and unchanged if you rescale or shift either variable, for example grams to kilograms.Eight students record hours studied (x) and exam score (y). Compute r step by step.
| student | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| hours x | 1 | 2 | 3 | 5 | 6 | 7 | 8 | 8 |
| score y | 48 | 52 | 55 | 58 | 62 | 66 | 69 | 70 |
| student | x − 5 | y − 60 | product | (x − 5)² | (y − 60)² |
|---|---|---|---|---|---|
| 1 | −4 | −12 | 48 | 16 | 144 |
| 2 | −3 | −8 | 24 | 9 | 64 |
| 3 | −2 | −5 | 10 | 4 | 25 |
| 4 | 0 | −2 | 0 | 0 | 4 |
| 5 | 1 | 2 | 2 | 1 | 4 |
So Sxy = 153, Sxx = 52, Syy = 458.
Notice what each part did. Student 4 studied 5 hours, exactly the average, so the x deviation is 0 and the product contributes nothing. Students 1 and 8 sit far from the centre on both axes and contribute the most. Correlation is dominated by the points that are far from both means.
Now let the code confirm every number, and move straight on to the case that breaks the rule.
Four datasets each have 11 points. All four have the same mean of x, the same mean of y, the same variance of x and y, the same r, and the same best-fit line. How alike do their scatter plots look?
Here are the four datasets, which are public and widely reproduced. Datasets I, II and III share the same x values. In set IV, ten points have x = 8 and one has x = 19.
| x (I, II, III) | 10 | 8 | 13 | 9 | 11 | 14 | 6 | 4 | 12 | 7 | 5 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| y I | 8.04 | 6.95 | 7.58 | 8.81 | 8.33 | 9.96 | 7.24 | 4.26 | 10.84 | 4.82 | 5.68 |
| y II | 9.14 | 8.14 | 8.74 | 8.77 | 9.26 | 8.10 |
| x (IV) | 8 | 8 | 8 | 8 | 8 | 8 | 8 | 19 | 8 | 8 | 8 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| y IV | 6.58 | 5.76 | 7.71 | 8.84 | 8.47 | 7.04 | 5.25 | 12.50 | 5.56 | 7.91 | 6.89 |
Running the code below confirms that for all four sets the mean of x is 9.00, the mean of y is 7.50, the variance of x is 11.00, the variance of y is about 4.13, the best-fit line is y ≈ 3.00 + 0.50x, and r is 0.816 for I, II and III and 0.817 for IV (about 0.816 in every case).
What the plots show:
Picture what r is built to detect. A scatter of points that forms a tilted oval cloud has r well away from zero, and the thinner the oval, the closer |r| is to 1. A round cloud has r near 0. A tilted oval is exactly the shape r can detect, and the only shape it can detect: a U or a ring is far from random, yet its r is near zero, which is Anscombe's set II again.
Spearman's ρ (rho) answers a looser question: when X goes up, does Y tend to go up, whatever the shape? The recipe is simple. Replace each value by its rank (1 for the smallest, n for the largest, ties share the average rank), then compute Pearson's r on the ranks.
| situation | Pearson r | Spearman ρ |
|---|---|---|
| straight line, light noise | good | about the same |
| monotone but curved (growth, saturation) | understates | captures it |
| a few extreme outliers | distorted | robust |
| non-monotone (U-shape) | near 0 | also near 0 |
| you need the actual linear slope | meaningful | not the right tool |
Spearman is not magic. A U-shape is not monotone, so it fails there too. And it discards the sizes of the gaps, so if magnitude matters, it throws information away.
In ML you read the matrix for three things:
y = 3·x1 + noise (so the true coefficients are 3 and 0). Refitting 300 times on fresh samples of 50, with x2 only loosely tied to x1 (r about 0.6), the coefficient on x1 has a standard deviation of 0.20. With x2 a near-copy (r about 0.998), the standard deviations balloon to about 2.85 each, a 14-fold jump, while the standard deviation of the sum of the two coefficients stays at 0.15 because the combined signal is still pinned down.0.05 to 0.5 and watch the coefficient spread fall as the two columns become less alike. The usual fixes for collinearity are to drop or combine redundant columns, or to add regularisation (RegularizationRegularizationRegularization adds a penalty term to the loss function (L1, L2) to discourage overly complex models and reduce overfitting.Learn more →), which adds a penalty on large coefficients and so shrinks the arbitrary coefficient splits towards a stable answer (covered in a later lesson).Two features have r = 0.998. You fit a linear regression with both. Which statement is most accurate?
df.corr() returns the Pearson correlation matrix of every numeric column at once (method="spearman" gives the rank version). df.groupby("col")[...].mean() splits the rows into groups and summarises each group, which is exactly the compare-inside-the-groups move that Simpson's paradox is about. And plt.imshow paints a grid of numbers as coloured squares, so a correlation matrix becomes a heat-map you can scan in a second: red for positive, blue for negative.users is the same data every time. We add one column, pro, which is 1 for a Pro user and 0 for a Free one, because a 0/1 column can be correlated like any other.churned against sessions_per_week is −0.561 and against session_min is −0.343: users who stay open the app more often and for longer. pro against churned is −0.290, so Pro users churn less. The grouped means say the same thing in the units you care about: stayers average 30.19 minutes and 4.35 sessions a week, churners 15.29 minutes and 1.20 sessions. And grouping by plan and device shows where the churn sits: Free Mobile churns 40.0 percent of its 70 users, Free Desktop 20.0 percent of 30, Pro Mobile 15.0 percent of 40 and Pro Desktop 6.7 percent of 60.Everything above is a correlation. It tells you that short sessions go with churn, not that lengthening a session would keep a user. Treat the heat-map as a list of places to look, and ask the three-explanations question of every red or blue square, which Part 2 spells out.
.corr() to .corr(method="spearman") and see which squares move. Then add df.groupby("plan")["session_min"].mean() and compare it with the grouped means above.Three short exercises put this lesson's main numbers in your hands: Spearman on a curved, tied dataset, the same correlations in pandas, and an outlier that moves Pearson but not Spearman. The pearson helper is given. Fill in the TODOs, then compare with the solution.
Tests · Pearson on the hours/steps data is about 0.758 and Spearman is about 0.994 (ties in steps get average ranks). df.corr() gives about 0.758 and df.corr(method='spearman') about 0.994 for hours against steps, and the mean of steps is 3.25 for late = 0 and 37.00 for late = 1. With the last steps value replaced by 1000, Pearson falls to about 0.596 while Spearman stays at about 0.994.
df.corr() repeats the Pearson number (0.758 between hours and steps) and method="spearman" gives 0.994. The late column, which marks the second half of the rows, correlates 0.873 with hours, and the grouped mean of steps rises from 3.25 to 37.00. In the third part, replacing the last value 95 by 1000 drags Pearson down from 0.7584 to 0.5958, while Spearman stays at 0.9940 because the ranks did not change.r = Sxy / √(Sxx · Syy); on the eight-student data Sxy = 153, Sxx = 52, Syy = 458, so r = 0.9914 and r squared = 0.983df.corr() gives the matrix, groupby().mean() the per-group view and imshow the heat-map; on StreamBox, churn correlates −0.561 with sessions per week and −0.343 with session minutesA dataset has Sxy = 60, Sxx = 25 and Syy = 196. What is Pearson's r?
| 6 | 2 | 6 | 12 | 4 | 36 |
| 7 | 3 | 9 | 27 | 9 | 81 |
| 8 | 3 | 10 | 30 | 9 | 100 |
| sum | 153 | 52 | 458 |
| 6.13 |
| 3.10 |
| 9.13 |
| 7.26 |
| 4.74 |
| y III | 7.46 | 6.77 | 12.74 | 7.11 | 7.81 | 8.84 | 6.08 | 5.39 | 8.15 | 6.42 | 5.73 |