Correlation, Causation & Simpson's Paradox
Here is where this lesson ends up. Three real-looking results that mislead, each of which you will be able to catch by the end:
- A classifier is trained on photos where most camels stand on sand and most cows on grass. It quietly learns "sand means camel", scores well on a test set from the same collection, and fails the day a camel stands on grass.
- A hospital model learns that patients given an intensive-care bed are higher risk. That is true in the data and useless as advice, because nobody should withhold the bed.
- Model A beats model B on overall accuracy while B is better inside every segment of users, which is Simpson's paradox in an evaluation table.
- Each is a correlation that is real in the data and still a poor guide to action. The sections below give you the questions that expose them.
After this lesson, you will be able to:
- Compute Pearson's r by hand from deviations, interpret r and r squared, and explain why r only detects straight-line association
- Use Anscombe's quartet to show that four very different datasets can share the same mean, variance, regression line and r of about 0.816
- Compute Spearman's rank correlation, and say when it beats Pearson (monotone but curved relationships, outliers)
- Read a correlation matrix, spot redundant features, explain why highly correlated inputs make linear regression coefficients unstable, and compute the matrix and per-group means in pandas
- Distinguish the three explanations for a correlation (X causes Y, Y causes X, Z causes both) and say what randomisation removes
- Work through a Simpson's paradox table from counts, and decide whether the pooled or the grouped table answers your question
Before You Start
#Why Correlation Is the Most Misused Number
Correlation answers one narrow question: if I know X, how well can I guess Y with a straight line? It says nothing about why, and nothing about what happens if you change X.
Cov(X, Y) = E[(X − μ_X)(Y − μ_Y)], where μ (mu) is the mean of a variable. This lesson takes the standardised version, Pearson's r, and puts it under stress. Causal estimation itself lives in the Causal Inference Basics lesson in the classical ML track, which covers propensity scores (the modelled chance that a unit receives the treatment, used to compare like with like) and Double ML (using machine-learning models to strip the confounders out of both variables before comparing them). Here we stay with the foundations: what the number is, how it breaks, and how to avoid being fooled by it.Every section below uses numbers you can verify by hand or by running the code. All counts and simulations are synthetic and labelled as such. The one real dataset is Anscombe's quartet, a famous public teaching example from 1973.
#Covariance Recap and Pearson r
Covariance measures whether X and Y sit above or below their means together:
The trouble with covariance is its units. Heights in centimetres and weights in kilograms give a covariance in cm·kg, and changing centimetres to metres changes the number by a factor of 100 without changing the relationship. Dividing by both standard deviations removes the units:
How to read the value:
- r = +1 means every point lies exactly on a rising line. r = −1 means exactly on a falling line.
- r = 0 means no linear association. It does not mean no association (see Anscombe below).
- r squared is the fraction of the variance of Y that the best straight line through X accounts for. With r = 0.5, only 25 percent.
- r is symmetric (
r(X, Y) = r(Y, X)) and unchanged if you rescale or shift either variable, for example grams to kilograms.
#Pearson r by Hand on Eight Points
Eight students record hours studied (x) and exam score (y). Compute r step by step.
| student | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| hours x | 1 | 2 | 3 | 5 | 6 | 7 | 8 | 8 |
| score y | 48 | 52 | 55 | 58 | 62 | 66 | 69 | 70 |
| student | x − 5 | y − 60 | product | (x − 5)² | (y − 60)² |
|---|---|---|---|---|---|
| 1 | −4 | −12 | 48 | 16 | 144 |
| 2 | −3 | −8 | 24 | 9 | 64 |
| 3 | −2 | −5 | 10 | 4 | 25 |
| 4 | 0 | −2 | 0 | 0 | 4 |
| 5 | 1 | 2 | 2 | 1 | 4 |
| 6 | 2 | 6 | 12 | 4 | 36 |
| 7 | 3 | 9 | 27 | 9 | 81 |
| 8 | 3 | 10 | 30 | 9 | 100 |
| sum | 153 | 52 | 458 |
So Sxy = 153, Sxx = 52, Syy = 458.
Notice what each part did. Student 4 studied 5 hours, exactly the average, so the x deviation is 0 and the product contributes nothing. Students 1 and 8 sit far from the centre on both axes and contribute the most. Correlation is dominated by the points that are far from both means.
Now let the code confirm every number, and move straight on to the case that breaks the rule.
#What r Measures and What It Does Not: Anscombe's Quartet
Four datasets each have 11 points. All four have the same mean of x, the same mean of y, the same variance of x and y, the same r, and the same best-fit line. How alike do their scatter plots look?
Here are the four datasets, which are public and widely reproduced. Datasets I, II and III share the same x values. In set IV, ten points have x = 8 and one has x = 19.
| x (I, II, III) | 10 | 8 | 13 | 9 | 11 | 14 | 6 | 4 | 12 | 7 | 5 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| y I | 8.04 | 6.95 | 7.58 | 8.81 | 8.33 | 9.96 | 7.24 | 4.26 | 10.84 | 4.82 | 5.68 |
| y II | 9.14 | 8.14 | 8.74 | 8.77 | 9.26 | 8.10 | 6.13 | 3.10 | 9.13 | 7.26 | 4.74 |
| y III | 7.46 | 6.77 | 12.74 | 7.11 | 7.81 | 8.84 | 6.08 | 5.39 | 8.15 | 6.42 | 5.73 |
| x (IV) | 8 | 8 | 8 | 8 | 8 | 8 | 8 | 19 | 8 | 8 | 8 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| y IV | 6.58 | 5.76 | 7.71 | 8.84 | 8.47 | 7.04 | 5.25 | 12.50 | 5.56 | 7.91 | 6.89 |
Running the code below confirms that for all four sets the mean of x is 9.00, the mean of y is 7.50, the variance of x is 11.00, the variance of y is about 4.13, the best-fit line is y ≈ 3.00 + 0.50x, and r is 0.816 for I, II and III and 0.817 for IV (about 0.816 in every case).
What the plots show:
- I: a noisy straight line. r is a fair summary here.
- II: a perfect smooth curve (an inverted U). The relationship is exact, yet r is still 0.816, because r only measures the straight-line part.
- III: an exact straight line with one outlier. Without the outlier r would be 1.000 and the line would be different. One point drags r down and tilts the slope.
- IV: every x is 8 except one point at 19. The entire correlation comes from that single point. Remove it and there is no relationship at all.
Picture what r is built to detect. A scatter of points that forms a tilted oval cloud has r well away from zero, and the thinner the oval, the closer |r| is to 1. A round cloud has r near 0. A tilted oval is exactly the shape r can detect, and the only shape it can detect: a U or a ring is far from random, yet its r is near zero, which is Anscombe's set II again.
#Spearman Rank Correlation
Spearman's ρ (rho) answers a looser question: when X goes up, does Y tend to go up, whatever the shape? The recipe is simple. Replace each value by its rank (1 for the smallest, n for the largest, ties share the average rank), then compute Pearson's r on the ranks.
| situation | Pearson r | Spearman ρ |
|---|---|---|
| straight line, light noise | good | about the same |
| monotone but curved (growth, saturation) | understates | captures it |
| a few extreme outliers | distorted | robust |
| non-monotone (U-shape) | near 0 | also near 0 |
| you need the actual linear slope | meaningful | not the right tool |
Spearman is not magic. A U-shape is not monotone, so it fails there too. And it discards the sizes of the gaps, so if magnitude matters, it throws information away.
#The Correlation Matrix and Multicollinearity
In ML you read the matrix for three things:
- Feature to label. A large |r| with the target suggests a feature carries a linear signal. Remember the limits above: it can also be a confounder's footprint.
- Feature to feature. Pairs near ±1 are redundant. Two columns that are the same quantity in different units (square metres and square feet) add no information, only trouble.
- Clusters of features. Blocks of high correlation hint at a smaller number of underlying factors, which is the idea behind PCA.
y = 3·x1 + noise (so the true coefficients are 3 and 0). Refitting 300 times on fresh samples of 50, with x2 only loosely tied to x1 (r about 0.6), the coefficient on x1 has a standard deviation of 0.20. With x2 a near-copy (r about 0.998), the standard deviations balloon to about 2.85 each, a 14-fold jump, while the standard deviation of the sum of the two coefficients stays at 0.15 because the combined signal is still pinned down.0.05 to 0.5 and watch the coefficient spread fall as the two columns become less alike. The usual fixes for collinearity are to drop or combine redundant columns, or to add regularisation (RegularizationRegularizationRegularization adds a penalty term to the loss function (L1, L2) to discourage overly complex models and reduce overfitting.Learn more →), which adds a penalty on large coefficients and so shrinks the arbitrary coefficient splits towards a stable answer (covered in a later lesson).Two features have r = 0.998. You fit a linear regression with both. Which statement is most accurate?
#The Same Tools in pandas: corr(), groupby and a Heat-Map
df.corr() returns the Pearson correlation matrix of every numeric column at once (method="spearman" gives the rank version). df.groupby("col")[...].mean() splits the rows into groups and summarises each group, which is exactly the compare-inside-the-groups move that Simpson's paradox is about. And plt.imshow paints a grid of numbers as coloured squares, so a correlation matrix becomes a heat-map you can scan in a second: red for positive, blue for negative.users is the same data every time. We add one column, pro, which is 1 for a Pro user and 0 for a Free one, because a 0/1 column can be correlated like any other.churned against sessions_per_week is −0.561 and against session_min is −0.343: users who stay open the app more often and for longer. pro against churned is −0.290, so Pro users churn less. The grouped means say the same thing in the units you care about: stayers average 30.19 minutes and 4.35 sessions a week, churners 15.29 minutes and 1.20 sessions. And grouping by plan and device shows where the churn sits: Free Mobile churns 40.0 percent of its 70 users, Free Desktop 20.0 percent of 30, Pro Mobile 15.0 percent of 40 and Pro Desktop 6.7 percent of 60.Everything above is a correlation. It tells you that short sessions go with churn, not that lengthening a session would keep a user. Treat the heat-map as a list of places to look, and ask the three-explanations question of every red or blue square.
.corr() to .corr(method="spearman") and see which squares move. Then add df.groupby("plan")["session_min"].mean() and compare it with the grouped means above.#Spurious Correlation and Confounders
Here is a confounder in a simulation. All numbers are synthetic. Let temperature Z drive both ice-cream sales X and sunburn cases Y over 500 days:
ice = 40 × temp + noiseburn = 3 × temp + noiseiceappears nowhere in the formula forburn. By construction, ice cream does not cause sunburn.
residual computes the leftovers.Running the simulation (seed 7) gives:
| quantity | value |
|---|---|
| r(temp, ice) | 0.928 |
| r(temp, burn) | 0.854 |
| r(ice, burn) | 0.784 |
| r(ice, burn) restricted to days between 24 and 26 degrees (102 days) | −0.032 |
| partial correlation (temperature regressed out of both) | −0.043 |
| r(ice, burn) after randomly shuffling which day gets which ice-cream figure | 0.032 |
A strong 0.784 correlation between ice cream and sunburn, with zero causal link. Hold temperature fixed (look at days of similar temperature) and the correlation vanishes. This is the signature of confounding: the association lives in the pooled data and disappears once you compare like with like.
#Correlation vs Causation: Three Explanations
When X and Y are correlated, at least one of these must be true (or several together):
- X causes Y. Studying more raises scores.
- Y causes X. Reverse causation. People with higher scores feel confident and study more, or sicker patients are given more medication so medication correlates with sickness.
- Z causes both. A common cause. Temperature raises both ice-cream sales and sunburn.
(A fourth, chance, is handled by sample size and the inference tools of the next lessons.) A correlation alone cannot tell the three apart. The scatter plot looks the same under all three.
The simulation shows the effect: shuffling the ice-cream column, which mimics assigning it at random, drops r from 0.784 to 0.032. The confounder's link is cut.
| tool | breaks which explanation | what it needs |
|---|---|---|
| randomisation (an A/B test) | Z causes both, and Y causes X | the ability to assign X |
| conditioning on a measured confounder | Z causes both | you measured Z, and Z is truly a common cause |
| time order (X before Y) | Y causes X | careful timing data |
| domain knowledge and mechanism | all of them | a credible causal story |
When you cannot randomise (you cannot assign people to smoke), you need the observational methods in the Causal Inference Basics lesson, and each one rests on assumptions you must argue for, since the data cannot verify them.
#Simpson's Paradox: When the Pooled Table Lies
| mild cases | severe cases | pooled | |
|---|---|---|---|
| Treatment A | 270 / 300 | 20 / 50 | 290 / 350 |
| Treatment B | 46 / 50 | 135 / 300 | 181 / 350 |
#Step 1: Compare inside the mild group
A recovered 270 of 300, which is 270 / 300 = 90.0 percent. B recovered 46 of 50, which is 46 / 50 = 92.0 percent.
#Step 2: Compare inside the severe group
A recovered 20 of 50, which is 20 / 50 = 40.0 percent. B recovered 135 of 300, which is 135 / 300 = 45.0 percent.
#Step 3: Pool the groups
A: (270 + 20) / (300 + 50) = 290 / 350 = 82.9 percent. B: (46 + 135) / (50 + 300) = 181 / 350 = 51.7 percent.
#Step 4: Find the reason
Look at who got which treatment. A treated 300 mild and only 50 severe patients. B treated 50 mild and 300 severe. A was handed mostly easy cases, which recover at high rates whatever you do. The pooled A rate is a weighted average, 300/350 = 86 percent weight on the easy group, while B's weights are 14 percent easy and 86 percent hard. The comparison was never like for like.
Drag the group-mix slider to change how many of A's patients are mild, and watch the pooled bars: find the mix where A's pooled lead flips to B, while B stays ahead inside both groups.
#Which table do you trust?
The arithmetic cannot tell you. The causal story does.
- Severity was decided before treatment and influences both who got which treatment and the outcome. Severity is a confounder. Use the grouped (or standardised) table. B is better.
- Suppose instead the grouping variable is something the treatment itself changes, for example "developed a complication", where treatment A causes complications that then lower recovery. Then splitting by that variable throws away part of the very effect you want to measure, and the pooled table is the honest one.
Same numbers, opposite advice. What changes is the direction of the arrows between treatment, the grouping variable and the outcome. That is why "just control for everything" is not a safe default.
In the synthetic table, treatment A looks better pooled but worse in both subgroups. What most directly produces this reversal?
#Simpson's paradox in a model's accuracy
pivot lays out accuracy by segment, and groupby(...).sum() pools the counts.Per segment, B wins both: 0.96 against 0.95 on easy requests and 0.40 against 0.30 on hard ones. Pooled, A wins by a mile, 885 / 1000 = 0.885 against 456 / 1000 = 0.456, only because A was handed mostly easy requests. On the same 50/50 mix, B scores 0.680 and A 0.625. Shipping A because of the pooled number would pick the worse model. The fix is the habit from the ML section below: report accuracy by segment, on the same segment mix, before declaring a model better.
#Selection Bias and Survivorship
Sometimes the correlation is created by how the data was collected, not by nature. Two common forms:
- Survivorship bias. You only see the units that survived a filter. If you study only funds that still exist after ten years, average returns look better than the full set of funds that were ever launched, because the poor performers closed and vanished from your data. If you study only the bullet holes on bombers that returned, you miss that planes hit elsewhere never came back.
- Selection into the dataset. Suppose a hiring model is trained only on people who were hired. Among hired people, test score and interview score can look negatively correlated even if they are positively correlated in the applicant pool, because anyone who was weak on both was rejected. Selecting on a variable that both depend on (here, being hired) manufactures a correlation. This is the collider effect described in the confounder section, arriving through the data pipeline: being hired is an effect of both scores, and you looked only at the hired.
A quick synthetic check you can run: draw 100,000 pairs of independent standard normals X and Y (seed 0, r = 0.002), keep only the rows where X + Y is above 1, and the kept rows (24,041 of them) show r of −0.61. Nothing about X and Y changed. Only the filter did.
#Connection to ML: Shortcuts and Observational Data
Almost all ML training data is observational, and almost all ML objectives reward correlation. A model minimises prediction error, and any feature that helps predict, causal or not, is fair game.
- Correlated features are not causal features. A feature with high |r| to the label is useful for prediction in the data you have. It is a safe lever only if it causes the label. A hospital risk model may learn that patients who got an intensive-care bed are higher risk, which is true in the data and useless as advice ("withhold the bed").
- Spurious shortcuts. A classifier can learn the background instead of the object if the two are confounded in training (for example, camels photographed on sand, cows on grass). It scores well on a test set drawn from the same collection, then fails when the background changes. The shortcut is a confounder that lives in the dataset, and the standard test split does not expose it because the test set has the same confounder.
- Distribution shift is a broken correlation. If a feature predicted the label only through a common cause, the prediction fails when that cause changes. Features whose relationship to the label is stable under change are the ones to prefer, which is one motivation for causal and invariant representation learning.
- Observational data can mislead a model about interventions. Predicting what will happen if we do X (a recommendation, a price change) is a different question from predicting what we will see given X. Models trained on logged data inherit the logging policy's confounding, which is why recommender teams run randomised experiments and use counterfactual evaluation.
- Simpson's paradox appears in metrics. An overall accuracy that rises while accuracy falls inside every user segment is entirely possible when the segment mix shifts, as the pandas table above showed (A pooled 0.885 against B 0.456, yet B better in both segments). Always report metrics by segment before declaring a model better.
For the causal tools that go beyond spotting the problem, see Causal Inference Basics in the classical ML track.
rng.normal(0, 80, n) to rng.normal(0, 800, n) so ice cream becomes much noisier. The pooled r drops, but the within-band r is still near zero. Noise weakens a confounded correlation without changing the fact that it is confounded.#Try It Yourself
Three short exercises put the lesson's main numbers in your hands: Spearman on a curved, tied dataset, a Simpson reversal from counts, and a partial correlation that makes a confounded link disappear. The pearson helper is given. Fill in the TODOs, then compare with the solution.
Tests · Pearson on the hours/steps data is about 0.758 and Spearman is about 0.994 (ties in steps get average ranks). Plan Q beats Plan P in both the young and old groups (94 percent vs 90 percent, 45 percent vs 20 percent), Plan P beats Plan Q pooled (76.0 percent vs 54.8 percent), and with a 50/50 mix Q gets 69.5 percent against P's 55.0 percent. The cafes and pharmacies correlation is about 0.706 and the partial correlation with city size removed is about -0.001.
When you run the solution you should see Pearson 0.7584 against Spearman 0.9940: the steps column rises every time, so its ranks nearly match the hours' ranks, and the two tied 3s share rank 2.5. In the Simpson part, Q beats P inside both groups (94.0 against 90.0 percent young, 45.0 against 20.0 percent old), yet P wins pooled, 76.0 against 54.8 percent, because P treated 200 young and only 50 old patients. At a 50/50 mix Q scores 69.5 percent and P 55.0 percent. In the third part, cafes and pharmacies correlate at 0.706 although neither causes the other, and with city size regressed out the partial correlation is −0.001.
#Key Takeaways
- r is covariance without units —
r = Sxy / √(Sxx · Syy); on the eight-student data Sxy = 153, Sxx = 52, Syy = 458, so r = 0.9914 and r squared = 0.983 - r only sees straight lines — Anscombe's four datasets share mean 7.50, variance 4.13, line y = 3 + 0.5x and r about 0.816, yet they are a line, a curve, a line with an outlier, and one lone point creating the correlation
- Spearman is Pearson on ranks — it gives 1.000 on a monotone curve where Pearson gives 0.7905, and 0.9286 with an outlier where Pearson gives 0.6955; it still fails on U-shapes
- Correlated inputs destabilise coefficients — with r about 0.998 between two features each coefficient's standard deviation rises from about 0.2 to about 2.85, while their sum stays stable at about 0.15
- pandas does the bookkeeping —
df.corr()gives the matrix,groupby().mean()the per-group view andimshowthe heat-map; on StreamBox, churn correlates −0.561 with sessions per week and −0.343 with session minutes - Three explanations, and randomisation removes two — X causes Y, Y causes X, Z causes both; assigning X at random (r from 0.784 to 0.032 in the simulation) cuts the common-cause and reverse-cause routes
- Simpson's paradox is unequal group mix — in the synthetic table B wins mild (92.0 vs 90.0) and severe (45.0 vs 40.0), yet A wins pooled (82.9 vs 51.7); whether to trust the pooled or grouped table depends on the causal story, not the arithmetic
- Metrics can reverse too — in the synthetic model table A pools at 0.885 against B's 0.456 while B wins both segments (0.96 vs 0.95 and 0.40 vs 0.30), so compare models by segment
- Models learn correlation, so audit the data — ask who is missing, test on data where the suspected shortcut is broken, and report metrics per segment
#Quick Check
A dataset has Sxy = 60, Sxx = 25 and Syy = 196. What is Pearson's r?