Probability & Bayes' Theorem
After this lesson, you will be able to:
- Understand conditional probability as narrowing your focus — 'given that X already happened, how likely is Y?'
- Apply Bayes' theorem to flip a probability around: go from 'how likely is the evidence if I am right' to 'how likely am I right given the evidence'
- Count people instead of juggling fractions: turn a test's two error rates into a table of 10,000 people and read Bayes' theorem off it
- Spot the base rate trap: understand why a test that catches 99% of sick people but also flags 5% of healthy people still gives mostly false alarms when the condition is rare
- Chain multiple evidence updates together using Bayes' theorem sequentially, as a spam filter does with each suspicious word it encounters
Before You Start
#Updating Your Beliefs
Probability is how smart people deal with uncertainty. You already use it every day — checking whether to carry an umbrella, deciding if a text is real or spam, guessing if your crush likes you back. This lesson just turns your gut feelings into precise numbers.
Probability is the language of uncertainty, and uncertainty is everywhere in AI. A classifier does not say "this is a cat" — it says "there is a 94% probability this is a cat." Understanding probability is not optional in machine learning. It is foundational.
#The Basics
A quick recap, because the previous lesson, Counting & Probability Foundations, built this up properly.
- A probability P(A) is a number from 0 (impossible) to 1 (certain) that says how likely event A is. The sample space is the list of every possible outcome, and the probabilities of all outcomes add up to 1.
- Read a probability as a long-run frequency: "70% chance of rain" means that in situations like this one, it rained about 7 times out of 10.
- Every classifier ends by turning raw scores into probabilities that add up to 1.
F.softmax(logits, dim=-1)is exactly that operation.
#The Basics: Quick Check
A spam filter assigns probability 0.92 to an email being spam. What does the number 0.92 mean?
#Conditional Probability
#Bayes' Theorem
P(rain | clouds) ≈ 0.40 in your city. What can you say about P(clouds | rain)?
The formula has four pieces, and each has a name:
- Posterior, written P(H|E): what we want, the probability of the hypothesis after seeing the evidence
- Likelihood, written P(E|H): how probable is this evidence if the hypothesis is true?
- Prior, written P(H): what we believed before seeing any evidence
- Evidence, written P(E): the total chance of seeing this evidence at all, whether or not the hypothesis is true
Rather than meet the formula cold, we will first count real people in a medical-test example, and then watch the formula appear as the same count written in symbols.
#The Medical Test: A Mind-Blowing Result
This is the single most counterintuitive result in probability. Understanding it deeply will change how you think about classification in ML.
A medical test catches 99% of sick people (sensitivity 99%) but also flags 5% of healthy people (false-positive rate 5%). The disease affects 1% of the population. You test positive. What is the probability you actually have the disease?
Let us count, with no formula at all. Imagine 10,000 people take the test. With a 1% prevalence (the share of people who have the disease), 100 are sick and 9,900 are healthy.
| Group | People | Test positive | Test negative |
|---|---|---|---|
| Sick (1%) | 100 | 99 caught | 1 missed |
| Healthy (99%) | 9,900 | 495 false alarms | 9,405 correctly cleared |
| Total | 10,000 | 594 | 9,406 |
The 99 comes from 99% of 100. The 495 comes from 5% of 9,900. Now answer the question by reading the table: you tested positive, so you are one of the 594 people in the positive column, and only 99 of those 594 are sick. That is 99 / 594 = 0.1667, about 16.7%.
Divide the top and bottom of that fraction by 10,000 and the table becomes Bayes' theorem. The 99 out of 10,000 is likelihood times prior (0.99 × 0.01). The 594 out of 10,000 is the total chance of a positive result, which is the evidence term P(E):
P(E) is "the total chance of a positive result": the sick people who test positive plus the healthy people who test positive. Let us rebuild it step by step.
#Step 1: The Prior (Base Rate)
The disease is rare. It affects 1 in 100 people.
P(Disease) = 0.01, P(No Disease) = 0.99
#Step 2: The Test's Two Error Rates (Likelihood)
The test catches 99% of sick people: if you have the disease, it correctly says "positive" 99% of the time.
P(Positive | Disease) = 0.99 (sensitivity)
But it also flags 5% of healthy people: if you are healthy, it incorrectly says "positive" 5% of the time.
P(Positive | No Disease) = 0.05 (false-positive rate)
#Step 3: Total Probability of Testing Positive
How often does the test say "positive" across the entire population? This is P(E), the evidence term, and it is the 594 out of 10,000 in the table.
P(Positive) = P(Pos|Disease) x P(Disease) + P(Pos|No Disease) x P(No Disease)
About 6% of everyone who gets tested will test positive.
#Step 4: Apply Bayes' Theorem
P(Disease | Positive) = P(Positive | Disease) x P(Disease) / P(Positive)
P(Disease | Positive) = 0.99 x 0.01 / 0.0594 = 0.0099 / 0.0594
#Step 5: The Shocking Answer
Why? The 100 sick people produce 99 true positives. But the 9,900 healthy people, 5% of whom are flagged, produce 495 false positives. So of 594 positive results, only 99 are truly sick: 99/594 = 16.7%.
#Visualizing the Base Rate Fallacy
The next picture draws all 10,000 people as dots, coloured by the four table cells. Drag the prevalence slider and watch the false alarms stop dominating, then drag the false-positive-rate slider and watch them shrink, and switch to the second tab to see the same answer as one area divided by two areas.
#Code the Bayesian update yourself
Reading the math is one thing — running it is another. The cell below codes the disease-test calculation from scratch and lets you sweep prevalence to see the posterior climb from "almost always a false alarm" to "almost certainly sick." A Monty Hall solver follows as a second snippet.
Monty Hall first, since its likelihoods are a big leap if they just appear. There are three doors and one car. You pick door 1. Monty knows where the car is, always opens a door you did not pick, and never opens the car. He opens door 3 and shows a goat. The likelihood is "how likely is Monty to open door 3, for each place the car could be?"
| Car is behind | Prior | What Monty can do | Chance he opens door 3 |
|---|---|---|---|
| Door 1 (your pick) | 1/3 | Door 2 and door 3 both hide goats, so he picks one at random | 1/2 |
| Door 2 | 1/3 | Door 3 is his only goat door to open (door 1 is yours, door 2 hides the car) | 1 |
| Door 3 | 1/3 | He cannot open the door with the car | 0 |
Multiply prior by likelihood: door 1 gets 1/3 × 1/2 = 1/6, door 2 gets 1/3 × 1 = 1/3, door 3 gets 1/3 × 0 = 0. These add to 1/2, so dividing each by 1/2 gives posteriors of 1/3, 2/3 and 0. Door 2 is twice as likely as door 1, so switching doubles your chance of winning.
When you run it, the first block prints P(positive) = 0.0594 and P(disease | positive) = 0.1667, and the sweep climbs from 0.019 at prevalence 0.001 to 0.952 at prevalence 0.50. The Monty Hall block prints 0.333, 0.667 and 0.000 for doors 1, 2 and 3, and the ratio of door 2 to door 1 is 2.0.
#Interactive: Watch Beliefs Update
The medical test was a single update. Beliefs can also be updated many times in a row, one piece of evidence at a time. The viz below tracks a different question: how often does this coin land heads? It starts with a flat prior (no opinion), and each flip is a new piece of evidence. The curve it draws is a Beta distribution, a flexible bump over the numbers from 0 to 1; a tall narrow bump means you are sure, a wide flat one means you are not (the Distribution Zoo lesson covers it properly).
Press the flip buttons and watch what the curve does as heads and tails accumulate, then press Flip 10 a few times and watch it narrow.
#Why Bayes Matters for ML
#Try It Yourself
A spam filter applies the update once per word, and each posterior becomes the next prior. Your job is to build that loop. Start from a 30% spam prior. The word "lottery" appears in 80% of spam and 1% of legitimate email, "free" in 60% of spam and 10% of legitimate email, and "meeting" in 2% of spam and 30% of legitimate email.
Tests · Verify the posteriors after lottery, free and meeting are 0.9717, 0.9952 and 0.9320, that the reversed order gives the same 0.9320, that unsubscribe raises it to 0.9910, and that the meeting likelihood ratio is 1 to 15.
Running the solution prints these values. After "lottery" the probability of spam is 0.9717, after "free" it is 0.9952, and after "meeting" it falls back to 0.9320. The reversed order ends at the same 0.9320, so the order of the words does not matter. Adding "unsubscribe" (0.40 versus 0.05) pushes it up to 0.9910, because that word is eight times more common in spam than in legitimate email. The ratio 0.02 / 0.30 is about 1 to 15, so "meeting" is evidence against spam, and it pulls a 99.5% belief back to 93.2%.
An Essay towards solving a Problem in the Doctrine of Chances
Thomas Bayes (posthumous) (1763)
The original paper by Bayes, published after his death by Richard Price. Remarkably accessible for an 18th-century mathematical paper. Introduces what we now call Bayes' theorem through a thought experiment involving billiard balls.
⚡ Playground: Probability Explorer → — adjust priors and likelihoods and watch posterior beliefs update live.
#Key Takeaways
- Bayes' theorem reverses conditional probabilities. It converts P(evidence|hypothesis) into P(hypothesis|evidence), letting you update beliefs as new data arrives
- The base rate fallacy is critical for ML. A test with 99% sensitivity and a 5% false-positive rate yields only a 16.7% chance of disease after a positive result when the condition is rare (1% prevalence), because false positives from the large healthy population overwhelm true positives
- Both levers matter. With sensitivity at 99%, cutting the false-positive rate from 5% to 1% lifts the answer to 50%, and raising prevalence from 1% to 3% lifts it to 38%
- Priors encode what you already know. Your initial belief (prior) gets updated by evidence (likelihood) to produce an updated belief (posterior); with enough evidence, the prior becomes irrelevant
- P(A|B) is not the same as P(B|A). This asymmetry is the entire reason Bayes' theorem exists and is a common source of reasoning errors in both medicine and ML
- Bayesian thinking pervades ML. From spam filters and Bayesian optimization to regularization (priors on weights) and generative models, Bayes' theorem provides the mathematical framework for reasoning under uncertainty
#Quick Check
A test catches 99% of sick people and has a 5% false-positive rate, and the disease has 1% prevalence (a positive test means a 16.7% chance of disease). Holding sensitivity at 99%, which change gives the higher chance of disease after a positive test: (a) the false-positive rate drops from 5% to 1%, or (b) the prevalence rises from 1% to 3%? Compute both.