You tested a new button. 7 clicks in 40 views. Is the click-through rate 17.5%? Maybe 10%? Maybe 30%? A single number hides the real answer, which is a whole range of plausible rates and how much you believe each. Bayesian inference gives you that range directly, updates it with every new view, and answers the question your boss actually asks: "what is the chance the new version is better?"
Learning Objectives
After this lesson, you will be able to:
Apply Bayes' rule to a whole parameter, not a single event: prior, likelihood, posterior over every possible value of θ
Do the Beta-Binomial update by hand (add clicks to α, add non-clicks to β) and read off the posterior mean, mode and 95% credible interval
Explain a prior as pseudo-counts, and see data overwhelm the prior as n grows
State the correct interpretation of a credible interval versus a confidence interval, and compute P(B beats A) with Monte Carlo
Know what to do when no closed form exists: grid approximation first, MCMC as the general tool
In the Bayes' theorem lesson you updated the probability of a single event: "is this email spam?" or "does this patient have the disease?" There the unknown was a yes/no fact. Here the unknown is a number: the true click-through rate θ of a button, somewhere between 0 and 1.
The recipe does not change. You still write posterior ∝ likelihood × prior. The only difference is that you now do it for every candidate value of θ at once, so the prior and posterior are curves, not single probabilities.
Two words in that formula need care. The likelihood is the probability of the data you saw, read as a function of the unknown rate. It is not a probability distribution over the rate. To see the difference, fix the data at 7 clicks in 40 views and slide the rate. If the true rate were 0.10, the chance of exactly 7 clicks in 40 views is 0.0576. At 0.175 it is 0.1640, and at 0.40 it is 0.0015. Those numbers score how well each rate explains the data. They do not add up to 1 across rates, so they are not beliefs about the rate until you multiply by a prior and rescale.
The sign ∝ means "equal up to a constant that does not depend on the rate". Constants such as the binomial coefficient and p(D) are the same for every candidate rate, so you can ignore them while you multiply, then rescale at the end so the curve has total area 1.
The posterior hands you the entire curve, so you can ask for a mean, a mode, an interval, or the probability that θ exceeds some threshold. The peak of the curve is called the MAP estimate (maximum a posteriori). A later lesson, Maximum Likelihood, takes the other route and keeps only the single best value.
Explore the single-event version first and notice how the belief moves with each observation. The parameter version is the same motion, over a continuum of values.
Try it: flip coins and watch the belief updateInteractive
Before the update rule, you need the curve it works on. A rate such as a click-through rate is a number between 0 and 1. The Beta distribution Beta(a, b) is a family of curves drawn over that range. The two positive numbers a and b set the shape. Think of a as a count of clicks and b as a count of non-clicks you already have in mind. A later section makes that precise.
Three curves are worth knowing by sight.
Beta(1, 1) is a flat line. Every rate from 0 to 1 is equally plausible before any data. That is not "no opinion": saying all rates are equally likely is itself a choice, and it says a 90% click rate is as plausible as 10%.
Beta(2, 8) is a hill leaning toward low rates. Its peak is at 0.125, its average is 0.2, and the middle 95% of the curve lies between 0.028 and 0.482, so it is still wide.
Beta(20, 80) has the same average, 0.2, but is far narrower: the middle 95% lies between 0.128 and 0.283. Larger a + b means a more confident belief.
The average of Beta(a, b) is a divided by (a + b). Pick a prior preset in the explorer below, leave the data at zero first to see the prior's shape, then open the Manual tab: it holds the example used next, where Beta(2, 8) plus 7 clicks in 40 views becomes Beta(9, 41). The second tab, StreamBox churn, uses the track's example streaming service, StreamBox. A user churns when they cancel, and the tab compares users on the Free plan with users on the Pro plan. You can ignore that tab for now.
You need a prior that lives on [0, 1] and a likelihood that counts clicks. The standard pairing is the Beta curve you just met for the prior and the Binomial for the data. The Binomial is the pattern of how many clicks you get in n views when each view clicks independently with the same probability θ. Their product has a remarkable property: it is another Beta.
θ∼Beta(α,β),k∣θ∼Binomial(n,θ)⟹θ∣k∼Beta(α+k,β+n−k)
Why does this work? Multiply the two curves and the exponents add. The Beta(α, β) density is proportional to θ^(α−1) (1−θ)^(β−1), and the binomial likelihood is proportional to θ^k (1−θ)^(n−k). With the numbers used below, the prior Beta(2, 8) is proportional to θ^1 (1−θ)^7, and the likelihood of 7 clicks in 40 views is proportional to θ^7 (1−θ)^33. The product is θ^(1+7) (1−θ)^(7+33) = θ^8 (1−θ)^40. A Beta(α + k, β + n − k) curve has exponents α + k − 1 and β + n − k − 1, so θ^8 (1−θ)^40 is exactly Beta(9, 41). When the prior and likelihood combine into a distribution of the same family as the prior, the prior is called conjugate.
Now the worked example. A button has prior Beta(2, 8): prior mean 0.2, equivalent to a history of 10 views with 2 clicks. You then observe 7 clicks in 40 views. Run it:
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The plot overlays three curves: the prior, the likelihood of the data (rescaled so it can share the axis) and the posterior. The posterior sits between the prior and the likelihood, a little closer to the likelihood because 40 real views outweigh 10 imaginary ones. You should see the posterior is Beta(9, 41) with mean 0.18, mode 0.1667, and a 95% credible interval of (0.0876, 0.2966). The MLE alone said 0.175. The mean is pulled slightly toward the prior mean 0.2, the mode (MAP) is a touch lower because Beta is skewed right, and the posterior also tells you that P(rate > 0.10) is about 0.948, a statement an MLE cannot make.
Data rarely arrive in one lump. Your 40 views came in over four days: 10 views each. Before you run anything, decide what you expect.
What Do You Think?
Start from Beta(2, 8). Day 1: 2 clicks in 10 views. Day 2: 3 in 10. Day 3: 1 in 10. Day 4: 1 in 10. You update the posterior after each day, using yesterday's posterior as today's prior. How does the final posterior compare with doing ONE update on all 40 views (7 clicks)?
Run it to see the intermediate beliefs:
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The printed beliefs are Beta(4, 16), Beta(7, 23), Beta(8, 32) and finally Beta(9, 41), with means 0.2000, 0.2333, 0.2000 and 0.1800. The single big update gives the same Beta(9, 41).
This is why Bayesian updating suits streaming data. You keep two running numbers, α and β, and never need to store the raw views. Yesterday's posterior is today's prior.
Beta(α,β)k1 clicks in n1Beta(α+k1,β+n1−k1)k2 clicks in n2Beta(α+k1+k2,β+(n1−k1)+(n2−k2))
#Prior Strength: Pseudo-Counts and Data Overwhelming
Read the update rule once more: α + k and β + n − k. The prior parameters sit exactly where counts of real data sit. So Beta(2, 8) behaves as if you had already watched 10 views and seen 2 clicks. That is the pseudo-count reading of a prior, and it makes the strength of a prior concrete: α + β is its effective sample size.
The posterior mean is a weighted average of the prior mean and the data rate, with weights proportional to the pseudo-count and the real count:
α+β+nα+k=wα+β+nα+β⋅α+βα+(1−w)⋅nk
Before the numbers, test your intuition on a prior that is much stronger than the data.
What Do You Think?
Now use a strong prior Beta(40, 160), which is 200 pseudo-views with prior mean 0.2. You observe the same data: 7 clicks in 40 views (rate 0.175). Where does the posterior mean land?
Now watch the prior fade. Hold the observed rate at 17.5% and let n grow:
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The table tells the story. At n = 0 the interval is the prior itself, (0.028, 0.482), width 0.454. At n = 40 the width is 0.209 and the prior's weight is 0.200. At n = 400 the width is 0.073 and the weight is 0.024. At n = 4000 the interval is (0.163, 0.187) with a prior weight of 0.002, so two people with different reasonable priors would now agree to three decimals. The strong prior from the quiz gives mean 0.1958 and interval (0.148, 0.248).
You can replay all of this in the Beta explorer from earlier in the lesson: choose the Strong preset, add the same 7 clicks in 40 views, and watch how little the curve moves.
In the statistical inference lesson you built a 95% confidence interval. The Bayesian counterpart is a 95% credible interval. They often look numerically similar, and they mean different things. On the same data, 7 clicks in 40 views, the code below builds two textbook confidence intervals. The Wald interval is the simple recipe from the Statistical Inference lesson: the estimate plus or minus 1.96 standard errors. Clopper-Pearson is a more cautious interval built from the exact Binomial distribution instead of a bell-curve approximation. The last lines measure coverage, which is the share of repeated experiments in which an interval contains the true value. A 95% interval should have coverage near 95%.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The results: the Wald confidence interval is (0.0572, 0.2928), the Clopper-Pearson interval is (0.0734, 0.3278), the flat-prior credible interval is (0.0882, 0.3206), and the Beta(2, 8)-prior credible interval is (0.0876, 0.2966). The four are close. The sentences you are allowed to say about them are not.
The coverage line in the output shows something else: at a true rate of 0.2 with n = 40, the Wald interval captured the truth only about 90.3% of the time, not 95%, and the flat-prior credible interval about 93.0%. Small samples and rates near zero are where textbook intervals wobble.
Quick check
A product manager asks: 'So is there a 95% chance the true click-through rate is between 0.0876 and 0.2966?' You computed that as a 95% credible interval. What is the correct reply?
Your team tests two checkout buttons. Variant A: 40 clicks in 400 views (10.0%). Variant B: 58 clicks in 420 views (13.8%). Is B better, or lucky? Give each variant a flat Beta(1, 1) prior, update to get two posteriors, then ask the direct question: what fraction of the time is a draw from B's posterior bigger than a draw from A's? That is Monte Carlo: draw many random samples and count.
Reading the output: P(B > A) = 0.9530. The median relative lift is 0.375 (a 37.5% improvement), with a 95% interval from −0.053 to 1.013, so the lift is probably large but a small loss is still possible. If you shipped B and were wrong, the expected loss is only about 0.00045 in click-through-rate units, which is 0.045 percentage points of click rate, tiny compared with the expected gain.
The frequentist test on the same counts gives z = 1.681, a two-sided p-value of 0.0928 and a one-sided p of 0.0464. At the conventional α = 0.05 the two-sided test says "not significant" while the Bayesian answer says "95% chance B is better". The two are closer than they look. With flat priors, P(B > A) = 0.953 is approximately 1 minus the one-sided p-value: 1 − 0.0464 = 0.9536. The p-value of 0.0928 is two-sided, because it counts a gap in either direction and so is twice the one-sided value. Most of the apparent disagreement is one-sided versus two-sided, not a different philosophy.
What still differs is the meaning. The one-sided p-value says how surprising a gap this large in B's favour would be if the buttons were identical. The posterior says how probable it is that B is better, given the data and the prior.
Does a Bayesian test let you peek for free? In one specific sense, yes. The posterior after the data you have seen is a valid statement about the rate, whether you planned 800 views or stopped because the numbers looked good, because the likelihood of the observed data does not depend on why you stopped. In another sense, no. A decision rule such as "ship as soon as P(B > A) passes 0.95" or "stop when the expected loss falls below a threshold" is a stopping rule, and a stopping rule changes how often the decisions it triggers are wrong. The Power, Sample Size & Multiple Testing lesson (section Peeking and Optional Stopping) showed this for p-values. The same thing happens here.
Beta-Binomial worked because the prior and likelihood fit each other. Several other pairs do too. You do not need to derive them here, just recognise the pattern: the posterior is the same family as the prior, with updated parameters.
Data model
Conjugate prior
What the update does
Binomial / Bernoulli (clicks)
Beta
Add successes to α, failures to β
Poisson (counts of events per hour or per day)
Gamma (a curve over positive numbers, such as an event rate)
Add total count to shape, add exposure to rate
Normal with known variance (mean unknown)
Normal
Precision-weighted average of prior mean and sample mean
Categorical / multinomial (one of several classes)
Dirichlet (the Beta extended to several classes: a curve over proportions that add to 1)
Add each class count to its parameter
When your prior does not fit the likelihood, there is no formula. Suppose you believe the button's rate is probably near 0.10 (like the sibling buttons) but could be near 0.30 (like the old design). That is a two-humped prior, a mixture of two Normals, and it is not conjugate to the Binomial.
The first fix is a grid approximation. Lay down a fine grid of candidate θ values, compute prior × likelihood at each grid point, and normalise by the sum. It is the general posterior-from-scratch recipe and it works whenever the parameter has one or two dimensions.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The grid gives posterior mean 0.1551, MAP 0.1285 and a 95% interval of (0.079, 0.298). The sanity check matters: with a flat prior the grid mean is 0.1905, matching the exact Beta(8, 34) mean 0.1905 to four decimals. Whenever you can check a grid against a closed form, do.
Grids collapse as dimensions grow: 100 grid points per parameter means 100^10 points for a ten-parameter model. The general tool is Markov Chain Monte Carlo (MCMC), which draws samples from the posterior without ever computing the normalising constant. Libraries such as PyMC, Stan and NumPyro implement it. That is a lesson of its own, and everything you did here (prior, likelihood, posterior, credible interval from samples) carries over unchanged.
A Naive Bayes spam filter estimates P(word | spam) from counts. If "meeting" never appeared in 100 spam words, the plain frequency is 0, and one zero wipes out the whole product of probabilities. Adding 1 to every count (Laplace smoothing) is exactly the posterior mean under a flat Dirichlet prior. With a vocabulary of 3 words: "free" goes from 0.1200 to 0.1262, "lottery" from 0.0500 to 0.0583, and the unseen "meeting" from 0.0000 to 0.0097. The unseen word is no longer impossible, just rare.
Suppose you have three buttons and every view is a chance to learn or to earn. Thompson sampling keeps a Beta posterior per button, draws one random rate from each, and shows the button with the biggest draw. Buttons that look good get chosen often; uncertain buttons still get tried because their wide posteriors sometimes produce a large draw. Exploration and exploitation fall out of the posterior with no extra tuning knob.
Tuning a learning rate or a prompt is expensive. Bayesian optimisation fits a posterior over the unknown score function (often a Gaussian process, a flexible curve-fitting model that reports its own uncertainty), then runs the experiment where the posterior says the gain is most likely or the uncertainty is largest. It is the same idea as Thompson sampling, applied to a continuous knob.
A point estimate says the next 100 views will bring about 17 or 18 clicks. The posterior predictive adds the uncertainty about θ itself on top of binomial noise, so its interval is honestly wider. That is the difference between a model that says "18" and a model that says "probably 8 to 30".
Run Thompson sampling against a naive uniform strategy. The three buttons have hidden true rates 0.05, 0.08 and 0.12; both strategies get 2000 views.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
Random assignment split views as [629, 675, 696], earned 182 clicks and had expected regret 71.0 (clicks lost compared with always showing the best button). Thompson sampling sent 1414 of 2000 views to the best button, earned 208 clicks, and cut regret to 30.3. Its estimates for the rarely shown buttons stay rough (0.074 and 0.089 against truths of 0.05 and 0.08) because it stopped spending views on them. That is the trade you chose: accuracy about the losers in exchange for earning more while learning. Results from a single seed are noisy, so try other seeds and see the pattern hold.
Finally, the posterior predictive. How many clicks should you expect in the next 100 views, given Beta(9, 41)?
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
The predictive mean is about 18.04 clicks with a 90% interval of (8, 30); plugging in the MLE gives (12, 24), a standard deviation of 3.79 versus 6.60. The plug-in forecast is overconfident because it treats an estimate as if it were the truth.
Run the whole Bayesian workflow on one landing-page test. The prior is Beta(2, 18), a guess of about 10% with the weight of 20 visits. Page A got 30 sign-ups in 300 visits and page B got 45 in 310. Work through the TODOs in order, and treat the last one as a stretch.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
Tests · Verify the posteriors are Beta(32, 288) and Beta(47, 283), the means are about 0.1000 and 0.1424, the intervals are about (0.0696, 0.1351) and (0.1069, 0.1821), P(B > A) is about 0.95 with seed 42, and in the stretch step most visits go to B with a small expected regret.
Your numbers should match these. A: posterior Beta(32, 288), mean 0.1000, 95% credible interval (0.0696, 0.1351). B: posterior Beta(47, 283), mean 0.1424, interval (0.1069, 0.1821). With seed 42, P(B > A) = 0.9518.
The two intervals overlap, yet the probability that B is better is about 95%. Overlap and "B beats A" are different questions, and the draws answer the second one directly.
In the stretch run, Thompson sampling sent 98.2% of the 2,000 visits to B, and the expected regret was 1.6 sign-ups against always showing B. That is the exploration the posterior does for you: once B's curve sits clearly higher, A almost stops being drawn. A different seed changes these two numbers a little, so try a few.
Beta(a, b) is a curve over a rate between 0 and 1: Beta(1, 1) is flat (a choice, not "no opinion"), Beta(2, 8) leans low with mean 0.2, and Beta(20, 80) has the same mean but is far narrower
The likelihood is the probability of the data you saw, read as a function of the unknown rate. It is not a distribution over the rate, and ∝ means equal up to a constant that does not depend on the rate
Bayesian inference for a parameter is Bayes' rule on a curve: posterior ∝ likelihood × prior over every candidate θ, giving a full distribution of belief rather than one number
The Beta-Binomial update is addition: Beta(α, β) plus k successes in n trials becomes Beta(α + k, β + n − k). Beta(2, 8) with 7 clicks in 40 views becomes Beta(9, 41) with mean 0.18, mode 0.1667 and 95% credible interval (0.0876, 0.2966)
Updating batch by batch gives exactly the same posterior as one big update, because yesterday's posterior is today's prior
A prior is pseudo-counts: α + β imaginary observations. Its weight is (α + β)/(α + β + n), so it fades as data grow, but a strong prior needs a lot of data to move
A credible interval is a probability statement about the parameter given the data and prior; a confidence interval is a statement about a long-run procedure. Do not read one as the other
A Bayesian A/B test samples both posteriors to get P(B > A) = 0.953 on data where the frequentist two-sided p-value was 0.093. With flat priors 0.953 is about 1 minus the one-sided p-value (0.0464), so most of the gap is one-sided versus two-sided. Report lift and expected loss with it
The posterior stays valid when you stop early, but a stopping rule such as "ship when P(B > A) passes 0.95" still changes how often the decisions it triggers are wrong
Without a closed form use a grid for one or two parameters and MCMC beyond that. In ML the same ideas appear as Laplace smoothing, Thompson sampling, Bayesian optimisation and predictive uncertainty
Beta(2, 8) and Beta(20, 80) both have mean 0.2. How do the two curves differ?
Next up: Information Theory & Entropy. You have measured how sure a posterior is; next you will measure how surprising data are, meet cross-entropy and KL divergence, and see why every classifier loss is a form of average surprise.
Your Reflection
Saves automatically
What’s one thing you learned? What’s still confusing?