The Distribution Zoo, Part 2: Fitting Data and Heavy Tails
A bell curve is a fine default, until you fit it to 60 API response times and it says that 2 of them finished in negative time. Real data is positive, skewed and has a tail, and the skill is to recognise the shape, pick a candidate and check it with numbers rather than by eye.
Where Part 1 left off: in The Distribution Zoo, Part 1: The Core Distributions you saw that every loss function hides a distribution, and you met the Uniform, Exponential, Beta, Dirichlet, Categorical, Multinomial and Student-t, each with a story, a mean and a variance. Beta updating is addition, the Exponential is memoryless, and a small sample needs the t multiplier. This part adds the skewed and heavy-tailed animals, then teaches the four tools that test whether a choice was right.
Learning Objectives
After this lesson, you will be able to:
Describe the generating story, the parameters, the mean and the variance of the Lognormal, Gamma and Chi-square distributions, and recognise a power law by its tail
Draw samples from each distribution with a seeded generator and check the sample mean and variance against the formulas, noting where heavy tails make the check loose
Explain why heavy tails make the sample mean unstable and why the Central Limit Theorem needs a finite variance
Score a model with log-likelihood and AIC, and compare it with the data using a Q-Q plot, a KS test and a Shapiro-Wilk test
Choose a candidate distribution from the data type with the decision guide, then check the choice with summary statistics, a scipy fit, AIC and a Q-Q plot
Turn a fitted distribution into an alert threshold from a percentile instead of mean plus 3 standard deviations
Many real quantities are built by multiplying, not adding. A request passes three stages, and each stage scales its time by a random factor, say 1.2, 0.9 and 1.5. The total scaling is 1.2 × 0.9 × 1.5 = 1.62. Take logs and the product becomes a sum: ln 1.2 + ln 0.9 + ln 1.5 = 0.182 − 0.105 + 0.405 = 0.482, and ln 1.62 = 0.482.
Sums of many small independent effects look bell-shaped (the Central Limit Theorem from the Descriptive Statistics lesson). So the logarithm of such a product is roughly Normal, and the quantity itself is Lognormal: positive, right-skewed, with a mean bigger than its median.
Parameters. μ and σ, the mean and standard deviation of the logarithm. The median is e^μ and the mean is e^(μ + σ²/2).
lnX∼N(μ,σ2),median=eμ,E[X]=eμ+σ2/2
Number example. With median 120 and log-scale standard deviation 0.6, the mean is 120 × e^(0.6²/2) = 143.7. Latencies, file sizes and session lengths are often lognormal.
Everything above is a formula until you run it. The cell below draws 10,000 samples from each main distribution with a seeded generator and prints the sample mean and variance next to the theoretical values from Part 1 and this lesson.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
What you should see. The first five rows agree with the formulas to within about 2%: for example the Beta(2, 5) sample mean is 0.2876 against 2/7 = 0.2857, and the Multinomial count has sample variance 2.4930 against the true 2.5. The last two rows are looser. The Student-t variance comes out 1.6436 against 1.6667, and the Lognormal variance 10,095 against 8,944, about 13% high. Heavy tails make the variance estimate jumpy: a few giant draws change it a lot. That is the practical meaning of a fat tail, and it comes back in the optional section below.
#Optional on a First Read: Gamma, Chi-Square and Heavy Tails
These three are the heavier or rarer members of the zoo. Nothing in the main path depends on them, so you can skip straight to The Decision Guide and return later. The Statistical Inference lesson touches the Chi-square again briefly, which is a good moment to come back.
Story. Add up k independent Exponential gaps and you get the total time until the k-th event. That sum is Gamma-distributed. It is also the flexible generalisation that fits any positive skewed quantity whose shape you want to tune.
Parameters. Shape k (how many events) and scale θ (the mean gap, which is 1/λ). Written Gamma(k, θ). With k = 1 you recover the Exponential.
E[T]=kθ,Var(T)=kθ2
Number example. Tickets arrive at 2 per minute (θ = 0.5 minutes). The time until the third ticket is Gamma(3, 0.5): mean 3 × 0.5 = 1.5 minutes, variance 3 × 0.25 = 0.75, standard deviation 0.866. The probability the third ticket has arrived by minute 1.5 is 0.5768. You can get the same number from the Poisson side: the expected count in 1.5 minutes is 2 × 1.5 = 3, and P(at least 3 tickets) = 0.5768. This duality is why Gamma and Poisson are partners: the Gamma asks "how long until k events?", the Poisson asks "how many events in this long?"
ML appearance. The Gamma is the conjugate prior for the rate of a Poisson (and for the precision, 1/variance, of a Gaussian), which means the same Beta-style update-by-addition works for count models. Chi-square, next, is a special case of it. Gamma regression also models positive, right-skewed targets such as insurance claim sizes.
Story. Take k independent standard normals Z₁, …, Z_k, square each, and add them up. The total is Chi-square with k degrees of freedom. Squared errors are exactly this kind of quantity, which is why it shows up whenever you ask how much variation is "too much".
Parameters. Degrees of freedom k. It is a Gamma special case (shape k/2, scale 2).
X=i=1∑kZi2∼χk2,E[X]=k,Var(X)=2k
Number example. Three standard normal draws (1.2, −0.5, 2.0) give 1.44 + 0.25 + 4 = 5.69. A χ² with 3 degrees of freedom has mean 3 and variance 6, and P(X > 5.69) = 0.128, so this is an unremarkable sum. The 5% cut-off for 3 degrees of freedom is 7.815.
A question it answers. A supplier claims a machine cuts parts with a spread of σ = 2 mm. You measure n = 10 parts and find a sample standard deviation s = 3 mm, so s² = 9. Is the machine really noisier than claimed, or is that luck?
Why is Chi-square the right tool? The sample variance s² is the sum of squared distances from the sample mean, divided by n − 1. Divide each distance by σ and every piece has the size of a standard normal, so (n − 1)s²/σ² is a sum of squared standard normals. It has n − 1 = 9 degrees of freedom rather than 10, because distances measured from the sample mean are forced to add to zero, which uses one up. If σ really is 2, this statistic averages 9.
Here the statistic is 9 × 9 / 4 = 20.25. For 9 degrees of freedom the middle 95% runs from 2.70 to 19.02, so 20.25 lands above it. The chance of a value this high when σ = 2 is 0.016. The data is evidence that the machine is noisier than claimed.
ML appearance. The Chi-square test checks whether observed class counts match expected counts (feature selection for categorical features, drift detection on category frequencies). The Chi-square also underlies the F-test for comparing models, and the sum of squared standardised residuals is the loss you minimise in least squares.
A distribution is heavy-tailed when extreme values are far more common than a Gaussian would allow. You have met one heavy-tailed story already, the Lognormal. Two more produce them.
Power law (Pareto, Zipf). "The rich get richer": the probability of exceeding x falls like x^(−α), not exponentially. For a Pareto with minimum 1 and α = 2, P(X > 10) = 0.01 and P(X > 100) = 0.0001. Multiplying x by 10 only divides the tail by 100, whereas a Normal would drop it astronomically. The mean exists only if α > 1 (for α = 2 it is 2), and the variance is infinite for α ≤ 2. Zipf's law is the same idea for ranks: the r-th most common word has frequency proportional to 1/r. In an idealised 1,000-word vocabulary with exponent 1, the top word takes 13.4% of all tokens, the second 6.7%, and the tenth only 1.3%. A handful of words dominate and a long tail of rare ones makes up the rest, which is why tokenizers, embeddings and vocabularies are designed around rare-word handling.
Mixtures and bursts. Real traffic mixes fast cached requests with slow cold ones, and bursts of activity with long silences. Mixtures of well-behaved distributions can produce tails that no single Normal can match.
#The Decision Guide: Which Distribution Fits My Data?
Start from what the numbers are, not what they look like. Then check the candidate against the data. The first column says how much each row matters on a first read: Core rows are used again later in this track, Optional rows can wait.
Matters most?
Your data
Candidate
Key parameters
Where it shows up in ML
Core
Yes/no outcome per trial
Bernoulli
p
Binary classification labels, dropout masks
Core
Count of successes in n trials
Binomial
n, p
Accuracy confidence intervals, A/B test conversions
Core
Count of events per interval
Poisson
λ
Click and error counts, count regression
Core
Time between events
Exponential
λ (mean 1/λ)
Queueing, survival analysis, simulation
Optional
Time until the k-th event; positive and skewed
Gamma
k, θ
Poisson-rate priors, claim-size regression
Five questions that narrow it down in under a minute.
Is the value discrete or continuous? Counts and classes use PMF-style distributions; measurements use densities.
What are its bounds? Below by zero suggests Exponential, Gamma or Lognormal. Between 0 and 1 suggests Beta. Unbounded on both sides suggests Normal or Student-t.
Is it symmetric? Compare mean and median. A mean well above the median signals right skew.
How fat is the tail? Compare the largest value to the mean in standard deviations, and look at the Q-Q plot.
What does the standard deviation say? If it equals the mean, think Exponential. If the data is positive with skew and standard deviation around half the mean, think Lognormal or Gamma.
Quick check
You are modelling how many seconds users stay on a page. The data is positive, strongly right-skewed, the median is 40 s, the mean is 65 s, and the standard deviation is 90 s. What is the best first guess?
The fit below uses four checking tools. Here is each one with a five-number example.
Log-likelihood. Score a model by the probability (density) it gives to each data point, take the log of each, and add them up. Closer to zero is better. The loss table in Part 1 is built from its negative.
AIC. More parameters always fit at least as well, so we charge a penalty: AIC = 2k − 2 × (log-likelihood), where k is the number of fitted parameters, and lower is better. Made-up example: model A has 2 parameters and log-likelihood −100, so AIC = 4 + 200 = 204. Model B has 1 parameter and log-likelihood −103, so AIC = 2 + 206 = 208. Model A wins by 4. As a rule of thumb, a gap of 2 or less is a toss-up and a gap in the tens is not.
Q-Q plot. It compares sorted data with the values a model expects at the same ranks. Take the five data points −1.2, −0.3, 0.1, 0.4, 1.9 and a standard Normal. At probabilities 0.1, 0.3, 0.5, 0.7, 0.9 the Normal's quantiles are −1.28, −0.52, 0, 0.52, 1.28. Pair them up: (−1.28, −1.2), (−0.52, −0.3), (0, 0.1), (0.52, 0.4), (1.28, 1.9). The first four sit close to the diagonal where data equals model. The last sits well above it: the data has a longer right tail than the model.
KS test (Kolmogorov-Smirnov). It measures the largest vertical gap between the data's cumulative fraction and the model's CDF, and reports a p-value: how often a gap this big would show up by luck if the model were true. A small p-value means the model is unlikely (the Statistical Inference lesson explains p-values properly). Example: the data 0.1, 0.2, 0.3, 0.4, 0.5 against Uniform(0, 1). At 0.5 the data's cumulative fraction is 1.0 while the model's CDF is 0.5, a gap of 0.5, and scipy gives p = 0.11. With only five points even a big gap is not strong evidence.
Shapiro-Wilk test. A test built for one question, "is this Normal?", with the same kind of p-value. On the skewed ten values 1, 1, 1, 2, 2, 3, 5, 9, 20, 40 it gives p = 0.0003, so Normal is very unlikely.
Now apply the guide. The cell below builds 60 synthetic response times (milliseconds), generated from a seeded lognormal so that everyone gets identical numbers. We pretend not to know the source, summarise the data, fit three candidates with scipy, rank them, and draw a histogram with fitted curves plus Q-Q plots.
What you should see.
Summary: mean 121.6 ms, median 113.8 ms, standard deviation 68.4 ms, skewness 1.60. The mean is above the median and skew is positive, so the data is right-skewed.
Normal fails: a Normal fitted to the data (mean 121.6, standard deviation 67.8) assigns probability 0.036 to negative response times, which means 2.2 of the 60 requests "expected" to take less than zero time. The Shapiro-Wilk test rejects normality with p of about 0.00002, and the Q-Q plot bends away from the line at both ends.
Exponential is the wrong skewed shape: its fingerprint is standard deviation equal to mean, but the data has 68.4 against 121.6. It also puts the most density at zero, while real response times have a floor. Its AIC is about 698.
Lognormal fits: the scipy fit gives a log-scale spread (sigma) of 0.529 and a median (scale) of 105.8 ms. The logarithm of the data passes Shapiro-Wilk with p of about 0.98, the KS test gives p of about 0.89, and the AIC is about 657 against 680 for the Normal.
A Pyodide-backed scratchpad for math lessons.
Loading visualization...
How to read the Q-Q plots. Each dot compares one sorted data value with the value the fitted model says should sit at the same rank. If the model is right, the dots hug the dashed line. A curve that bends upward at the right means the real tail is heavier than the model's; points below zero on the model axis (Normal case) mean the model allows impossible values. The lognormal panel is nearly straight; the Normal panel curls at both ends.
Make the decision. The log-scale fit says: median 105.8 ms, spread parameter 0.529. Use it to set an alert threshold from a percentile instead of "mean plus 3 standard deviations": the fitted p95 is 252.7 ms, against 233.1 ms from the Normal fit and an observed 268 ms in this sample. The Normal underestimates the tail you care about, and it is the tail that pages you at night.
Tests · The summary line shows n=40, mean 86.0, median 70.5 and skew 1.88. The ratio std/mean is 0.69, far from the 1 an Exponential needs. AIC is 443.3 for the Normal and 415.4 for the lognormal, and the Normal puts 0.0716 of its mass below 0 ms. Shapiro-Wilk gives p of about 4.3e-06 on the raw data and 0.351 on the log data. The fitted p95 is 182.6 for the Normal and 185.4 for the lognormal, against an observed 203.1.
When you run the solution you should see n=40, mean 86.0, median 70.5 and skew 1.88, so the data is right-skewed. The ratio std/mean is 0.69, far from the 1 that an Exponential needs. AIC is 443.3 for the Normal and 415.4 for the lognormal, a gap of 27.9 in favour of the lognormal, and the Normal puts 0.0716 of its mass below 0 ms. Shapiro-Wilk gives p = 4.3e-06 on the raw data and 0.351 on the log data, so the log is plausibly Normal. The fitted p95 is 182.6 ms for the Normal and 185.4 ms for the lognormal, against an observed 203.1 ms. With 40 points the observed p95 rests on the top few values, so both fits sit a little low, and the AIC gap is the stronger evidence.
Lognormal is what multiplication produces — take logs and a product becomes a sum; with median 120 and log-scale spread 0.6 the mean is 143.7, which is why you report percentiles for skewed data
Sampling checks the formulas — 10,000 seeded draws land within a couple of percent of the theoretical mean and variance, except where the tail is heavy, as with the Lognormal variance, 13% high
Gamma, Chi-square and power laws extend the zoo — the third ticket at 2 per minute is Gamma(3, 0.5) with mean 1.5 minutes; a Chi-square statistic of 20.25 on 9 degrees of freedom flagged a noisy machine; a Pareto with α = 2 has infinite variance, so its sample mean keeps jumping
Four tools check a fit — AIC = 2k − 2 × log-likelihood (lower wins, a gap of 2 or less is a toss-up), a Q-Q plot shows where a model goes wrong, and KS and Shapiro-Wilk give p-values
The decision guide starts with the data type — discrete or continuous, bounded or not, symmetric or skewed, fat tail or thin, then verify the pick with a fit; in the worked example a lognormal beat the Normal by about 23 AIC points and the Exponential by about 41
Model A has 3 fitted parameters and a log-likelihood of −200. Model B has 2 fitted parameters and a log-likelihood of −204. Which one does AIC prefer?
Next up: Sampling, Standard Error & the Central Limit Theorem. You will see why an average of samples wobbles, how much it wobbles, and why that wobble turns bell-shaped as the sample grows.