The Distribution Zoo: Which Distribution Fits My Data?
A bell curve is a fine default, until you model response times with it and predict that 2 in 60 requests finish in negative time. Every distribution is a story about how data got generated: waiting for a bus, flipping a weighted coin, adding up squared errors, multiplying small random effects. Learn the stories and you can look at a column of numbers and say which distribution it probably came from.
Learning Objectives
After this lesson, you will be able to:
Explain how the loss function you train with (squared error, absolute error, cross-entropy) is the negative log of a distribution you assumed for the target
Describe the generating story, the parameters, the mean and the variance of the Uniform, Exponential, Beta, Dirichlet, Categorical, Multinomial, Student-t and Lognormal distributions
Use the memoryless property and the Poisson link to turn event counts into waiting times, and the other way round
Read a Beta distribution as a tally of successes and failures and update it with data, then generalise the same idea to k classes with a Dirichlet
Draw samples from each distribution with a seeded generator and check the sample mean and variance against the formulas
Choose a candidate distribution from the data type, then check the choice with summary statistics, a scipy fit, AIC and a Q-Q plot
Choosing a distribution is a modelling decision, not a formality. Pick the shape that matches how the data was generated and everything downstream (loss functions, confidence intervals, anomaly thresholds) gets better for free.
You have already met three of the animals. In the Descriptive Statistics lesson you saw the Normal, the Bernoulli, the Binomial and the Poisson. In the Random Variables lesson you learned that a PMF gives probabilities for countable values, a PDF gives density for continuous ones, and a CDF accumulates them. This lesson does not repeat those. It adds the most useful of the rest of the zoo and, more importantly, teaches you to choose.
#The Payoff First: Every Loss Function Hides a Distribution
Here is the reason this lesson matters for machine learning. When a model predicts and reality answers, you need one number for how bad the prediction was. The principled number is the negative log of the probability the model gave to what actually happened. Add it up over the data and you have the negative log-likelihood, and training minimises it. Three tiny cases show what comes out.
Yes or no. A classifier says a click has probability 0.8, and a click happens. The loss is −ln 0.8 = 0.223. Had it said 0.1, the loss would be −ln 0.1 = 2.303, ten times worse. That is binary cross-entropy, and the story behind it is the Bernoulli.
A number with Normal noise. A regression model predicts 1, the truth is 3, and the noise has standard deviation 1. The negative log of the Normal density at 3 is ½ × (3 − 1)² + 0.919 = 2.919. The 0.919 never changes with the prediction, so training only feels the squared error. That is mean squared error, and the story is the Normal.
A number with Laplace noise. The Laplace distribution has a sharper peak than the Normal and fatter tails, and its density is ½e^(−|x − μ|) when its scale is 1. The same prediction gives −ln(½e^(−2)) = 2 + ln 2 = 2.693: the absolute error plus a constant. That is absolute error (L1), and the story is the Laplace.
If the target is
Assume this distribution
The loss you get
Yes or no
Bernoulli
Binary cross-entropy
One of k classes
Categorical
Cross-entropy on the softmax output
A count of events
Poisson
Poisson loss, λ − y ln λ (plus a constant)
A real number, symmetric noise
Normal
Squared error
A real number, spiky noise
Laplace
Absolute error
A real number with outliers
Student-t
A loss that grows slowly for big errors
The last row has numbers too. Move a prediction 20 standard units away from the truth. Squared error charges 200. A Student-t loss with 5 degrees of freedom and scale 1 charges about 13.2, so one wild point cannot dominate training. Priors work the same way: a Normal prior on the weights adds an L2 penalty, and a Laplace prior adds an L1 penalty (the MLE, MAP and Fisher Information lesson shows why).
So choosing a distribution is choosing a loss. The rest of the lesson is about choosing well.
Look at one data point and ask what it is. This short table is your first filter, and the full guide with parameters comes after you have met each story.
One data point looks like
Try this story
Yes or no
Bernoulli (add many up and you get a Binomial count)
A count: 0, 1, 2, … events in a window
Poisson
A positive time until something happens, with no memory
Exponential
A number between 0 and 1 that is itself a probability
Beta (one probability) or Dirichlet (several that sum to 1)
One label out of k
Categorical (counts of labels: Multinomial)
Any real number, errors symmetric
Normal (with outliers or very few samples: Student-t)
Every entry below follows the same five-part template: story (how the data gets made), parameters, mean and variance, a number you can check by hand, and where it shows up in ML.
Warm up with the Poisson, because the Exponential section turns it inside out. Set λ = 2 and read the bar heights: P(0) is about 0.135, P(1) and P(2) are both about 0.271, and P(3) is about 0.180.
Loading visualization...
Remember these four numbers. They return in the Exponential section, where the same λ = 2 describes the gaps between events instead of the counts.
Story. A number is drawn from an interval and no region is favoured. A random-number generator, a spinner, a shuffled index.
Parameters. Lower bound a and upper bound b. The density is flat: f(x) = 1/(b − a) on [a, b], and zero elsewhere.
E[X]=2a+b,Var(X)=12(b−a)2
Number example. Xavier (Glorot) initialisation for a layer with 256 inputs and 256 outputs (the fan-in and fan-out) draws weights from Uniform(−L, L) with L = √(6 / (256 + 256)) ≈ 0.1083. The variance is (2L)²/12 = 0.0039, which equals 2/(fan_in + fan_out) = 2/512. The width of the interval was chosen so that this variance keeps signal sizes stable from layer to layer.
ML appearance. Random initialisation, random sampling of mini-batches and hyper-parameter search all start from uniform draws. Uniform is also the engine behind every other sampler: draw U ~ Uniform(0, 1), push it through the inverse CDF of any distribution, and you get a sample from that distribution. The inverse CDF is the CDF run backwards: give it a probability such as 0.5 and it returns the value with that much probability below it, here the median. That trick is called inverse transform sampling, and you will use it on the Exponential next.
Story. Events arrive at a steady average rate, independently of each other: support tickets, packets, clicks. The Poisson distribution counts how many arrive in a fixed window. The Exponential measures the gap until the next one.
Parameters. Rate λ (events per unit time). Density f(t) = λe^(−λt) for t ≥ 0.
E[T]=λ1,Var(T)=λ21,P(T>t)=e−λt
Number example. Tickets arrive at λ = 2 per minute. The average gap is 0.5 minutes and the standard deviation is also 0.5. The chance of waiting more than a minute is e^(−2) ≈ 0.1353. Notice that this is exactly the Poisson P(0) you read off the slider above, and it must be: "no ticket in the next minute" and "the next gap is longer than a minute" are the same event.
What Do You Think?
Support tickets arrive randomly, and the gaps between them are Exponential with an average of 5 minutes. It has now been 20 minutes since the last ticket. How much longer do you expect to wait for the next one?
The memoryless property. For an Exponential, P(T > s + t | T > s) = P(T > t). Check it with λ = 2: P(T > 3 | T > 2) = e^(−6) / e^(−4) = e^(−2) ≈ 0.1353, the same as P(T > 1). The past waiting time carries no information about the future, because the underlying events are independent. It is the only continuous distribution with this property.
Sampling it. The Exponential's CDF is F(t) = 1 − e^(−λt). Set U = F(t) for U ~ Uniform(0, 1) and solve for t: T = −ln(1 − U)/λ, which is the inverse CDF from the Uniform section at work. With λ = 2 and U = 0.5 you get T = 0.3466 minutes, which is the median gap (ln 2 / λ). Half of all gaps are shorter than 0.35 minutes even though the mean is 0.5; the distribution is right-skewed.
ML appearance. Time-to-event models in survival analysis start from the Exponential, which has a constant hazard rate (the same chance of the event in the next instant, however long you have waited). Queueing models for model-serving capacity assume exponential inter-arrival times. In reinforcement learning, continuous-time MDPs use exponential holding times, and in simulation it is the default way to generate random event times.
Story. You want to describe your belief about an unknown probability p, such as a click rate. Every value of p lives in [0, 1], and you want the shape to sharpen as evidence arrives. The Beta distribution is exactly a density on [0, 1].
Parameters. Two positive numbers a and b. The cleanest way to read them: start from the flat curve Beta(1, 1), which means "nothing seen yet", then add 1 to a for every success and 1 to b for every failure. So Beta(a, b) behaves like a notebook holding a − 1 successes and b − 1 failures.
Two readings of the centre differ slightly, and it pays to keep them apart. The most likely value (the mode) is (a − 1)/(a + b − 2), which is exactly the raw success rate in the notebook. The mean a/(a + b) is pulled toward 0.5, because it behaves as if the notebook also held one extra success and one extra failure.
Number example. Start with Beta(2, 2): mean 0.5, standard deviation 0.224, which is the flat start plus one imagined success and one imagined failure, a mild belief that the rate is near the middle. Observe 8 clicks and 2 non-clicks. The posterior is Beta(2 + 8, 2 + 2) = Beta(10, 4), a notebook with 9 successes and 3 failures. Its mode is 9/12 = 0.75, its mean is 10/14 = 0.714, its standard deviation is 0.117, and its 95% interval runs from 0.462 to 0.909. The raw rate 8/10 = 0.8 got pulled toward the prior. The exercise at the end adds 50 more observations, and you will watch the interval tighten.
Updating beliefs this way is the whole subject of the Bayesian Inference in Practice lesson, which covers priors, credible intervals and A/B tests properly. Here you only need the shape, the mean, and the update-by-addition rule.
Use the explorer to confirm the formulas above. Choose Beta and watch the mean, variance and mode readouts as you move α and β.
Try it: switch families and move the slidersInteractive
Loading visualization...
Try this: Select Exponential and set λ = 2; the readout should show mean 0.5 and variance 0.25. Select Beta and set α = 2, β = 5 and confirm the mean is 2/7 = 0.286. Then set α = 4 and β = 10 (the slider maximum is 10) and compare with α = 2, β = 5: the mean stays at 0.286 but the curve tightens, which is what more evidence looks like. Select Uniform with a = 0, b = 1 and confirm the variance reads 0.083.
ML appearance. Thompson sampling for bandits keeps a Beta for every arm's success rate, samples one value from each, and plays the highest. The Beta prior is also how Bayesian A/B testing talks about conversion rates.
Quick check
You observe 3 successes and 1 failure with a flat Beta(1, 1) prior. What is the posterior and its mean?
Story. Instead of one probability you need a whole probability vector: three topic proportions, or the class probabilities of a classifier. Each entry is non-negative and they sum to 1. The set of all such vectors is called the simplex; for three classes it is a triangle whose corners are "all class A", "all class B" and "all class C". The Dirichlet is a distribution over points of the simplex. With k = 2 it is the Beta.
Parameters. One pseudo-count per class: α = (α₁, …, α_k), with α₀ = α₁ + … + α_k. As with the Beta, bigger numbers mean more evidence.
Where the formulas come from. Look at one entry p_i and lump all the other classes together as "not i". The Dirichlet then reduces to a Beta(α_i, α₀ − α_i), so the mean and variance above are just the Beta formulas with a = α_i and b = α₀ − α_i. The covariance comes from the sum-to-1 rule: if p₁ + p₂ + p₃ is always exactly 1, its variance is 0, and that forces the pairwise covariances to be negative.
Number example. Dirichlet(2, 3, 5) has α₀ = 10. The expected proportions are (0.2, 0.3, 0.5). The variance of the first entry is 2 × 8 / (100 × 11) = 0.01455 (standard deviation 0.121), and the covariance between the first two is −2 × 3 / (100 × 11) = −0.00545, negative as it must be. Check the sum-to-1 rule: the three variances (0.01455, 0.01909, 0.02273) add to 0.05636, the three covariances (−0.00545, −0.00909, −0.01364) add to −0.02818, and 0.05636 + 2 × (−0.02818) = 0 exactly. Multiply every α by 10 and the means are unchanged but the standard deviations shrink to about a third.
ML appearance. A softmax output is one point on the simplex, which is exactly where a Dirichlet lives. A Dirichlet is what you reach for when you want to say how unsure you are about the probabilities themselves. Latent Dirichlet Allocation (a topic model that treats each document as a mix of topics) draws each document's topic mix, and each topic's word mix, from Dirichlet priors. With all α below 1, samples concentrate near the corners: documents mostly about one topic.
Story. A Categorical distribution is one draw from k classes with probabilities p₁, …, p_k: a die roll, a predicted token, a softmax output sampled once. A Multinomial repeats that n times and records the count per class. Bernoulli and Binomial are the k = 2 cases.
Parameters. The probability vector p (summing to 1), and for the Multinomial also the number of draws n.
Number example. Ten draws from p = (0.5, 0.3, 0.2). The expected counts are (5, 3, 2). The probability of landing on exactly (5, 3, 2) is 10! / (5! 3! 2!) × 0.5⁵ × 0.3³ × 0.2² = 2520 × 0.03125 × 0.027 × 0.04 = 0.08505, about 8.5%. Even the single most likely outcome happens less than one time in ten. The softmax of logits (2, 1, 0.1) gives the categorical distribution (0.659, 0.242, 0.099).
ML appearance. The categorical is the output distribution of every classifier and of every language model (the next token is one categorical draw over the vocabulary). Cross-entropy loss is the negative log of the probability the categorical gave to the true class, as in the table at the top: with these probabilities, a true class 1 costs −ln 0.659 = 0.417 and a true class 2 costs 1.417. Bag-of-words models and Naive Bayes treat a document's word counts as one Multinomial draw.
Story. You estimate a mean from a small sample and must also estimate the spread from the same few points. That extra uncertainty makes the standardised mean fatter-tailed than a Normal. The Student-t is the distribution of (sample mean − true mean) / (s/√n) when the data is Normal, and it is also a model of its own for data with outliers.
Parameters. Degrees of freedom ν (for the t-test, ν = n − 1). As ν grows, it converges to the standard Normal. Small ν means thick tails. At ν = 1 it is the Cauchy distribution, which has no mean at all.
E[T]=0(ν>1),Var(T)=ν−2ν(ν>2)
Number example (tails). How likely is a value more than 4 standard units from zero? For a standard Normal it is 0.0000633, about 1 in 15,800. For a t with ν = 5 it is 0.0103, about 1 in 97. That is roughly 163 times more likely. The variance of t₅ is 5/3 = 1.667, so even the typical spread is 29% wider than a standard Normal's standard deviation of 1 (√1.667 = 1.29). Those fat tails are why the Student-t loss in the table at the top forgives big errors.
Number example (the t-test). The 95% two-sided multiplier is 1.960 for a Normal. For a t it depends on the sample size: 12.71 with n = 2, 4.30 with n = 3, 2.776 with n = 5, 2.262 with n = 10, 2.045 with n = 30. If you used 1.96 on samples of size 5, your "95%" intervals would miss the true mean about 12% of the time in simulation (they would cover only about 88%); the t multiplier brings coverage back to 95%. Run the cell below to see it yourself.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
ML appearance. The t-test (from the Statistical Inference lesson) is the Student-t at work: comparing two models over a handful of cross-validation folds. Student-t likelihoods give regression models robustness to outliers, and t-SNE uses a t distribution with ν = 1 in the low-dimensional space precisely because its heavy tails let clusters separate cleanly.
Many real quantities are built by multiplying, not adding. A request passes three stages, and each stage scales its time by a random factor, say 1.2, 0.9 and 1.5. The total scaling is 1.2 × 0.9 × 1.5 = 1.62. Take logs and the product becomes a sum: ln 1.2 + ln 0.9 + ln 1.5 = 0.182 − 0.105 + 0.405 = 0.482, and ln 1.62 = 0.482.
Sums of many small independent effects look bell-shaped (the Central Limit Theorem from the Descriptive Statistics lesson). So the logarithm of such a product is roughly Normal, and the quantity itself is Lognormal: positive, right-skewed, with a mean bigger than its median.
Parameters. μ and σ, the mean and standard deviation of the logarithm. The median is e^μ and the mean is e^(μ + σ²/2).
lnX∼N(μ,σ2),median=eμ,E[X]=eμ+σ2/2
Number example. With median 120 and log-scale standard deviation 0.6, the mean is 120 × e^(0.6²/2) = 143.7. Latencies, file sizes and session lengths are often lognormal.
Everything above is a formula until you run it. The cell below draws 10,000 samples from each main distribution with a seeded generator and prints the sample mean and variance next to the theoretical values from this lesson.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
What you should see. The first five rows agree with the formulas to within about 2%: for example the Beta(2, 5) sample mean is 0.2876 against 2/7 = 0.2857, and the Multinomial count has sample variance 2.4930 against the true 2.5. The last two rows are looser. The Student-t variance comes out 1.6436 against 1.6667, and the Lognormal variance 10,095 against 8,944, about 13% high. Heavy tails make the variance estimate jumpy: a few giant draws change it a lot. That is the practical meaning of a fat tail, and it comes back in the optional section below.
#Optional on a First Read: Gamma, Chi-Square and Heavy Tails
These three are the heavier or rarer members of the zoo. Nothing in the main path depends on them, so you can skip straight to The Decision Guide and return later. The Statistical Inference lesson touches the Chi-square again briefly, which is a good moment to come back.
Story. Add up k independent Exponential gaps and you get the total time until the k-th event. That sum is Gamma-distributed. It is also the flexible generalisation that fits any positive skewed quantity whose shape you want to tune.
Parameters. Shape k (how many events) and scale θ (the mean gap, which is 1/λ). Written Gamma(k, θ). With k = 1 you recover the Exponential.
E[T]=kθ,Var(T)=kθ2
Number example. Tickets arrive at 2 per minute (θ = 0.5 minutes). The time until the third ticket is Gamma(3, 0.5): mean 3 × 0.5 = 1.5 minutes, variance 3 × 0.25 = 0.75, standard deviation 0.866. The probability the third ticket has arrived by minute 1.5 is 0.5768. You can get the same number from the Poisson side: the expected count in 1.5 minutes is 2 × 1.5 = 3, and P(at least 3 tickets) = 0.5768. This duality is why Gamma and Poisson are partners: the Gamma asks "how long until k events?", the Poisson asks "how many events in this long?"
ML appearance. The Gamma is the conjugate prior for the rate of a Poisson (and for the precision, 1/variance, of a Gaussian), which means the same Beta-style update-by-addition works for count models. Chi-square, next, is a special case of it. Gamma regression also models positive, right-skewed targets such as insurance claim sizes.
Story. Take k independent standard normals Z₁, …, Z_k, square each, and add them up. The total is Chi-square with k degrees of freedom. Squared errors are exactly this kind of quantity, which is why it shows up whenever you ask how much variation is "too much".
Parameters. Degrees of freedom k. It is a Gamma special case (shape k/2, scale 2).
X=i=1∑kZi2∼χk2,E[X]=k,Var(X)=2k
Number example. Three standard normal draws (1.2, −0.5, 2.0) give 1.44 + 0.25 + 4 = 5.69. A χ² with 3 degrees of freedom has mean 3 and variance 6, and P(X > 5.69) = 0.128, so this is an unremarkable sum. The 5% cut-off for 3 degrees of freedom is 7.815.
A question it answers. A supplier claims a machine cuts parts with a spread of σ = 2 mm. You measure n = 10 parts and find a sample standard deviation s = 3 mm, so s² = 9. Is the machine really noisier than claimed, or is that luck?
Why is Chi-square the right tool? The sample variance s² is the sum of squared distances from the sample mean, divided by n − 1. Divide each distance by σ and every piece has the size of a standard normal, so (n − 1)s²/σ² is a sum of squared standard normals. It has n − 1 = 9 degrees of freedom rather than 10, because distances measured from the sample mean are forced to add to zero, which uses one up. If σ really is 2, this statistic averages 9.
Here the statistic is 9 × 9 / 4 = 20.25. For 9 degrees of freedom the middle 95% runs from 2.70 to 19.02, so 20.25 lands above it. The chance of a value this high when σ = 2 is 0.016. The data is evidence that the machine is noisier than claimed.
ML appearance. The Chi-square test checks whether observed class counts match expected counts (feature selection for categorical features, drift detection on category frequencies). The Chi-square also underlies the F-test for comparing models, and the sum of squared standardised residuals is the loss you minimise in least squares.
A distribution is heavy-tailed when extreme values are far more common than a Gaussian would allow. You have met one heavy-tailed story already, the Lognormal. Two more produce them.
Power law (Pareto, Zipf). "The rich get richer": the probability of exceeding x falls like x^(−α), not exponentially. For a Pareto with minimum 1 and α = 2, P(X > 10) = 0.01 and P(X > 100) = 0.0001. Multiplying x by 10 only divides the tail by 100, whereas a Normal would drop it astronomically. The mean exists only if α > 1 (for α = 2 it is 2), and the variance is infinite for α ≤ 2. Zipf's law is the same idea for ranks: the r-th most common word has frequency proportional to 1/r. In an idealised 1,000-word vocabulary with exponent 1, the top word takes 13.4% of all tokens, the second 6.7%, and the tenth only 1.3%. A handful of words dominate and a long tail of rare ones makes up the rest, which is why tokenizers, embeddings and vocabularies are designed around rare-word handling.
Mixtures and bursts. Real traffic mixes fast cached requests with slow cold ones, and bursts of activity with long silences. Mixtures of well-behaved distributions can produce tails that no single Normal can match.
#The Decision Guide: Which Distribution Fits My Data?
Start from what the numbers are, not what they look like. Then check the candidate against the data. The first column says how much each row matters on a first read: Core rows are used again later in this track, Optional rows can wait.
Matters most?
Your data
Candidate
Key parameters
Where it shows up in ML
Core
Yes/no outcome per trial
Bernoulli
p
Binary classification labels, dropout masks
Core
Count of successes in n trials
Binomial
n, p
Accuracy confidence intervals, A/B test conversions
Sum of squared standardised errors; variance tests
Chi-square
k
Goodness-of-fit, feature selection, F-tests
Core
Positive, right-skewed, from multiplied effects
Lognormal
μ, σ of the log
Latency, file sizes, session durations
Optional
Extreme tail, "a few giants and many tiny"
Pareto / Zipf
α
Word frequencies, user activity, gradient noise
Core
Bounded, no preference between values
Uniform
a, b
Initialisation, random search, sampling engine
Five questions that narrow it down in under a minute.
Is the value discrete or continuous? Counts and classes use PMF-style distributions; measurements use densities.
What are its bounds? Below by zero suggests Exponential, Gamma or Lognormal. Between 0 and 1 suggests Beta. Unbounded on both sides suggests Normal or Student-t.
Is it symmetric? Compare mean and median. A mean well above the median signals right skew.
How fat is the tail? Compare the largest value to the mean in standard deviations, and look at the Q-Q plot.
What does the standard deviation say? If it equals the mean, think Exponential. If the data is positive with skew and standard deviation around half the mean, think Lognormal or Gamma.
Quick check
You are modelling how many seconds users stay on a page. The data is positive, strongly right-skewed, the median is 40 s, the mean is 65 s, and the standard deviation is 90 s. What is the best first guess?
The fit below uses four checking tools. Here is each one with a five-number example.
Log-likelihood. Score a model by the probability (density) it gives to each data point, take the log of each, and add them up. Closer to zero is better. The loss from the top of this lesson is its negative.
AIC. More parameters always fit at least as well, so we charge a penalty: AIC = 2k − 2 × (log-likelihood), where k is the number of fitted parameters, and lower is better. Made-up example: model A has 2 parameters and log-likelihood −100, so AIC = 4 + 200 = 204. Model B has 1 parameter and log-likelihood −103, so AIC = 2 + 206 = 208. Model A wins by 4. As a rule of thumb, a gap of 2 or less is a toss-up and a gap in the tens is not.
Q-Q plot. It compares sorted data with the values a model expects at the same ranks. Take the five data points −1.2, −0.3, 0.1, 0.4, 1.9 and a standard Normal. At probabilities 0.1, 0.3, 0.5, 0.7, 0.9 the Normal's quantiles are −1.28, −0.52, 0, 0.52, 1.28. Pair them up: (−1.28, −1.2), (−0.52, −0.3), (0, 0.1), (0.52, 0.4), (1.28, 1.9). The first four sit close to the diagonal where data equals model. The last sits well above it: the data has a longer right tail than the model.
KS test (Kolmogorov-Smirnov). It measures the largest vertical gap between the data's cumulative fraction and the model's CDF, and reports a p-value: how often a gap this big would show up by luck if the model were true. A small p-value means the model is unlikely (the Statistical Inference lesson explains p-values properly). Example: the data 0.1, 0.2, 0.3, 0.4, 0.5 against Uniform(0, 1). At 0.5 the data's cumulative fraction is 1.0 while the model's CDF is 0.5, a gap of 0.5, and scipy gives p = 0.11. With only five points even a big gap is not strong evidence.
Shapiro-Wilk test. A test built for one question, "is this Normal?", with the same kind of p-value. On the skewed ten values 1, 1, 1, 2, 2, 3, 5, 9, 20, 40 it gives p = 0.0003, so Normal is very unlikely.
Now apply the guide. The cell below builds 60 synthetic response times (milliseconds), generated from a seeded lognormal so that everyone gets identical numbers. We pretend not to know the source, summarise the data, fit three candidates with scipy, rank them, and draw a histogram with fitted curves plus Q-Q plots.
What you should see.
Summary: mean 121.6 ms, median 113.8 ms, standard deviation 68.4 ms, skewness 1.60. The mean is above the median and skew is positive, so the data is right-skewed.
Normal fails: a Normal fitted to the data (mean 121.6, standard deviation 67.8) assigns probability 0.036 to negative response times, which means 2.2 of the 60 requests "expected" to take less than zero time. The Shapiro-Wilk test rejects normality with p of about 0.00002, and the Q-Q plot bends away from the line at both ends.
Exponential is the wrong skewed shape: its fingerprint is standard deviation equal to mean, but the data has 68.4 against 121.6. It also puts the most density at zero, while real response times have a floor. Its AIC is about 698.
Lognormal fits: the scipy fit gives a log-scale spread (sigma) of 0.529 and a median (scale) of 105.8 ms. The logarithm of the data passes Shapiro-Wilk with p of about 0.98, the KS test gives p of about 0.89, and the AIC is about 657 against 680 for the Normal.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
How to read the Q-Q plots. Each dot compares one sorted data value with the value the fitted model says should sit at the same rank. If the model is right, the dots hug the dashed line. A curve that bends upward at the right means the real tail is heavier than the model's; points below zero on the model axis (Normal case) mean the model allows impossible values. The lognormal panel is nearly straight; the Normal panel curls at both ends.
Make the decision. The log-scale fit says: median 105.8 ms, spread parameter 0.529. Use it to set an alert threshold from a percentile instead of "mean plus 3 standard deviations": the fitted p95 is 252.7 ms, against 233.1 ms from the Normal fit and an observed 268 ms in this sample. The Normal underestimates the tail you care about, and it is the tail that pages you at night.
Tests · After day 1 the posterior is Beta(10, 4) with mean 0.714. After day 2 it is Beta(50, 14) with mean 0.781 and sd 0.051, and the 95% interval 0.673 to 0.873 is narrower than after day 1. The mode is 0.790, closer to the raw rate 0.8 than the mean 0.781. The conditional exponential survival ratio equals sf(5), about 0.3679. The loss is 0.2231 at a predicted 0.8 and 2.3026 at a predicted 0.1.
When you run the solution you should see Beta(10, 4) with mean 0.714 after day 1, then Beta(50, 14) with mean 0.781 and standard deviation 0.051 after day 2. The 95% interval tightens from 0.462 to 0.909 down to 0.673 to 0.873: more evidence, tighter belief. The mode 0.790 sits closer to the raw rate 0.800 than the mean 0.781 does, as the notebook reading predicts. The memoryless ratio and sf(5) both print 0.3679. The loss is 0.2231 when the model said 0.8 and 2.3026 when it said 0.1.
Every distribution is a story, and your loss function already assumed one — Normal gives squared error, Laplace absolute error, Bernoulli binary cross-entropy, Categorical cross-entropy; Exponential is the gap between random events, Beta is a belief about a probability, Dirichlet is a belief about a probability vector
Exponential is memoryless and its standard deviation equals its mean — with rate 2 per minute the mean gap is 0.5 minutes, P(T > 1) = 0.1353, and that equals the Poisson probability of zero events
Beta parameters are a notebook of tallies and updating is addition — Beta(2, 2) plus 8 clicks and 2 misses becomes Beta(10, 4), mode 0.75 and mean 0.714; the Dirichlet does the same for k classes, and a softmax output is a point on its simplex
Small samples and outliers need the Student-t — with n = 5 the 95% multiplier is 2.776, not 1.96, and the t has variance ν/(ν − 2)
Sampling checks the formulas — 10,000 seeded draws land within a couple of percent of the theoretical mean and variance, except where the tail is heavy, as with the Lognormal variance, 13% high
The decision guide starts with the data type — discrete or continuous, bounded or not, symmetric or skewed, fat tail or thin, then verify the pick with a fit; in the worked example a lognormal beat the Normal by about 23 AIC points and the Exponential by about 41
You train a regression model with mean absolute error (L1) loss. Which noise story are you implicitly assuming for the target?
Next up: Sampling, Standard Error & the Bootstrap. You will see why an average of samples wobbles, how much it wobbles, and how to measure that wobble by resampling your own data.
Your Reflection
Saves automatically
What’s one thing you learned? What’s still confusing?