MLE, MAP & Fisher Information
loss = nn.CrossEntropyLoss() the same way again — you'll know what it's actually optimizing.After this lesson, you will be able to:
- Flip your perspective from probability to likelihood — the same number p(data | θ), viewed as a function of θ instead of data
- Derive Maximum Likelihood Estimators (MLE) from first principles for Bernoulli, categorical, and Gaussian models
- See that cross-entropy and MSE — the two losses that train every modern AI model — are negative log-likelihoods in disguise
- Understand Maximum A Posteriori (MAP) estimation as MLE plus a prior, and recognize L2 / L1 regularization as Gaussian / Laplace priors
Before You Start
#The Conceptual Flip: Probability vs Likelihood
MLE is the engine that powers every supervised learning algorithm. Once you see that cross-entropy and MSE are both negative log-likelihoods, the entire ML loss-function zoo collapses into a single principle: pick the model that makes the data most probable.
Throughout this lesson, θ (the Greek letter theta) is the unknown number we want to learn from data. For a coin, θ is the chance of heads.
Before we crank the algebra, lock in your gut on the probability-vs-likelihood split.
You flip a coin three times and observe HHH. Which statement is TRUE about the likelihood L(θ | HHH) = θ³ as a function of θ?
The likelihood is the same number as the probability — same formula, same arithmetic. The only thing that changes is which variable is fixed and which is free. In probability, θ is fixed (we know the coin is fair) and data varies (we ask "how likely are 3 heads?"). In likelihood, data is fixed (we saw 3 heads) and θ varies (we ask "which θ best explains this?").
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
#Maximum Likelihood Estimation
For independent identically distributed (iid) observations, the joint probability factorizes:
#Why log-likelihood?
Two reasons, both critical:
- Numerical stability: Computers represent floating-point numbers with finite precision. Multiplying 10,000 numbers each ≈ 0.001 gives 10^(-30,000), which is exactly zero in any practical floating-point format. Logging converts the product to a sum:
log Π p_i = Σ log p_i. Sums do not underflow. - Mathematical convenience: The log is monotonically increasing, so
argmax_θ f(θ) = argmax_θ log f(θ). The maximum stays the same. But the log-likelihood is a SUM, and derivatives of sums are sums of derivatives — much easier to differentiate term-by-term.
#Worked MLE: Bernoulli
This is the canonical example. Master it and the rest follow.
P(X=x | θ) = θ^x (1-θ)^(1-x) covers both cases.For n iid flips with k heads:
A note on shape. The log-likelihood ℓ(θ) is concave: a downward bowl with exactly one peak. The likelihood L(θ) itself, before the log, can look like a bell-shaped hump (for 3 heads in 4 flips it is a skewed hump peaking at 0.75). Both peak at the same θ, which is why taking the log never moves the answer.
k / θ = (n − k) / (1 − θ)
k (1 − θ) = (n − k) θ
k − kθ = nθ − kθ
k = nθ
θ̂_MLE = k / n
This is the simplest derivation in MLE-land but every other MLE follows the same recipe: write likelihood → take log → take derivative → set to zero → solve.
Let's compute the Bernoulli MLE numerically and watch it converge to the truth as n grows. The playground below draws n Bernoulli samples with θ=0.7, computes the MLE k/n, and plots the log-likelihood — drag the cursor to see how the peak narrows as evidence accumulates.
#How Sure Is the MLE? Curvature and Standard Error
Work it with real numbers. You flip a coin 500 times and see 150 heads, so θ̂ = 150 / 500 = 0.3. Curvature is how quickly the slope of ℓ changes, which is minus the second derivative. Differentiate the slope from Step 3 once more:
At θ̂ = 0.3 the first term is 150 / 0.09 = 1666.67 and the second is 350 / 0.49 = 714.29. Their sum is 2380.95. That number, the size of the bend, is the information I. At the peak k = nθ̂, so the two terms simplify to n / θ̂ + n / (1 − θ̂), which equals n / (θ̂ (1 − θ̂)). Check: 500 / (0.3 × 0.7) = 2380.95.
So θ̂ = 0.30 with a standard error of 0.0205, roughly 0.30 plus or minus 0.02. A rough 95% range is θ̂ plus or minus 2 SE, about 0.26 to 0.34. The count n sits in the numerator of I, so more flips sharpen the peak. With ten times as many flips and the same 30% heads (n = 5000, k = 1500), I = 23809.5 and SE = 0.00648, which is the old SE divided by √10.
Leave the overlay on in the picture below and drag n up: the curve for 10n flips is far narrower than the curve for n flips, and the table reads out the curvature and the standard error 1 over √information.
Now predict how the information depends on θ̂ before you open the second tab.
For n coin flips the information is I(θ) = n / [θ(1 − θ)]. At which θ is I LARGEST, so the estimate is most precise in absolute terms?
The fair coin is the hardest to pin down in absolute terms. I is smallest at θ = 0.5 (equal to 4n) and grows without bound toward 0 and 1, while the standard error is largest at 0.5. Open the Information vs theta tab of the picture above and slide θ to see the bowl shape of I.
One more check on what the information does when the sample size changes.
You observe HHHHT (4 heads, 1 tail). The Bernoulli MLE is θ̂ = 0.8. If you instead observed HHHHHHHHTT (8 heads, 2 tails), same empirical rate, what happens to the MLE and to the information (the SHARPNESS of the likelihood peak)?
#Worked MLE: Categorical (k classes)
Generalize from binary (Bernoulli) to k classes. Start with numbers. You label 100 emails: 10 promotions, 20 social, 70 primary. The natural guess for the three class probabilities is 0.10, 0.20 and 0.70. MLE agrees, and here is why.
ℓ(θ) = Σ_i n_i log θ_i. Maximize subject to Σ_i θ_i = 1.The MLE for a categorical is what every imbalanced-class baseline uses: predict the most frequent class with probability equal to its empirical frequency.
#Worked MLE: Gaussian
Here is the whole surface at once. Each point of the heatmap is a pair (μ, σ), and its brightness is the log-likelihood of the sample you draw. Change the sample size n and watch the bright region tighten around the star, which marks the MLE; leave the prior switch off for now, we use it in the MAP section.
∂ℓ/∂μ = 0. The first term has no μ. The second term gives (1/σ²) Σ (x_i − μ) = 0, so μ̂_MLE = (1/n) Σ x_i = x̄ (the sample mean).∂ℓ/∂v = −n / (2v) + S / (2v²) = 0
multiply both sides by 2v²: −n v + S = 0
v = S / n
Put in μ̂ = x̄ and you get the result below. Check with x = [1, 2, 3, 4, 5]: x̄ = 3, S = 4 + 1 + 0 + 1 + 4 = 10, so σ̂² = 10 / 5 = 2.0.
Drag mu and sigma until the curve hugs the points, then press Jump to the MLE and watch the negative log-likelihood reach its minimum at the sample mean.
#The Punchline: Cross-Entropy IS Negative Log-Likelihood
This is the moment that ties the entire math track together.
A k-class classifier outputs probabilities p = (p_1, ..., p_k). The true label is class y_true, encoded as a one-hot vector y where y_i = 1 if i = y_true and 0 otherwise.
The categorical likelihood for this one example is:
The negative log-likelihood is therefore:
F.cross_entropy(logits, targets) is doing Maximum Likelihood Estimation under a categorical likelihood. Image classifiers, language models and speech recognizers trained with cross-entropy are all MLE machines.When you read about "cross-entropy loss" in a paper, replace it mentally with "negative log-likelihood." It will feel less arbitrary and more inevitable.
Slide the predicted probability toward the wrong answer and watch cross-entropy climb while squared error flattens, then open the Softmax tab with logits 2, 1, 0.
#MSE IS Negative Log-Likelihood Under Gaussian Noise
Same trick, regression edition.
y = f_θ(x) + ε where the noise ε ~ N(0, σ²) (the standard regression assumption).The likelihood of observation y given input x and parameters θ is the Gaussian density centered at f_θ(x):
The negative log-likelihood across n iid examples is:
This implicit assumption matters. If your residuals are heavy-tailed (extreme values show up far more often than a bell curve allows, as with financial returns or latencies), Gaussian MLE is suboptimal. The right losses for those cases:
- Heavy tails / outliers → Huber loss = MLE under a "Gaussian in the middle, Laplace in the tails" noise model.
- Symmetric heavy tails → MAE (mean absolute error) = MLE under Laplace noise.
- Asymmetric → Quantile loss = MLE for asymmetric Laplace.
Every "robust regression loss" is just MLE under a different noise model.
#MAP: MLE Plus a Prior
Taking logs:
You have a Beta(2, 2) prior on θ (a 'weak' prior centred at 0.5). You see 3 heads in 4 flips. The MLE is 0.75. Where does the MAP land?
Run it numerically. The cell below computes both MLE and MAP for the same Bernoulli data, sweeps n, and shows that the gap between them collapses as n grows — that's the prior fading.
#L2 Regularization = Gaussian Prior MAP
The most common prior in ML: weights are a priori Gaussian-distributed around zero.
For Gaussian likelihood + Gaussian prior, the MAP objective becomes:
NLL(θ; data) + (λ/2) ‖θ‖² = MSE + L2 penalty = ridge regression
This is the classical derivation of ridge regression from Bayesian assumptions. Here λ (lambda) is the penalty strength, and it is the precision of the prior, which is 1 divided by its variance. Be precise about the scaling:
- If the loss is the SUM of the per-example negative log-likelihoods plus (λ/2)‖w‖², the prior on each weight is N(0, 1/λ). With λ = 100 the variance is 0.01 and the standard deviation is 0.1. With λ = 1 the standard deviation is 1.
- If the loss is the MEAN over n examples plus (λ/2)‖w‖², multiply through by n to get back to the sum form: the penalty becomes (nλ/2)‖w‖², so the prior precision is nλ and the variance is 1/(nλ). With n = 100 and λ = 4 the prior variance is 0.0025 and the standard deviation is 0.05, not 0.5. So the factor of n matters.
- If the loss is plain squared error instead of the full Gaussian NLL, the noise variance σ² is absorbed into λ as well.
weight_decay as a Gaussian prior is an analogy about the shrink-toward-zero effect, not an exact identity. Also keep the scale in mind: trained weights are small. A common initialization for large transformers, GPT-2's, draws weights with standard deviation 0.02, so a prior standard deviation of 10 would say almost nothing.#L1 Regularization = Laplace Prior MAP
|θ|/b for a scalar.LASSO regression, sparse coding, and L1-regularized neural networks all implement Laplace-prior MAP estimation.
You minimize (the SUM of per-example negative log-likelihoods) + (λ/2)‖w‖² with λ = 100. From a MAP perspective, what prior are you assuming on each weight?
#How MLE/MAP Power Modern AI
#Cross-entropy training in PyTorch
loss = F.cross_entropy(logits, targets) and call loss.backward(), you are computing the negative log-likelihood of the data under your model's predicted categorical distribution, then backpropagating its gradient. Image classification, language modeling and speech recognition trained this way are doing categorical MLE. The loss is not an arbitrary choice; it is the maximum-likelihood objective falling out of the categorical assumption.#Weight decay and L2 penalties as Gaussian-prior MAP
A loss of the form (sum of per-example NLL) + (λ/2)‖w‖² is MAP estimation under a zero-mean Gaussian prior with variance 1/λ. The prior keeps weights small unless the data forces them to be large, which is the point of regularization. AdamW's decoupled weight decay has the same shrink-toward-zero effect and is best read as an analogy to this prior, not an exact identity.
#Changing the noise model changes the loss
MSE assumes Gaussian noise, MAE assumes Laplace noise, Huber assumes Gaussian in the middle and Laplace in the tails. When your residuals have heavy tails, you do not search for a loss; you pick the noise model that matches your data and take its negative log-likelihood.
Theory of Statistical Estimation
Ronald A. Fisher (1925)
#Try It Yourself
You have met the likelihood, the MAP prior and Fisher information as formulas. Here you run all three on one small dataset: 500 coin flips from a coin whose true heads probability is 0.3. Work through the TODOs in order, then open the solution and compare.
Tests · Verify heads=138 of 500, NLL at 0.5 about 346.57 and at the MLE about 294.57, gradient descent p and mean(x) both 0.2760, prior mode 0.1250, MAP 0.2736 between the prior mode and the MLE 0.2760, Fisher I about 2502.2, SE about 0.0200 and bootstrap SD about 0.0201.
The seeded run gives 138 heads in 500 flips, so the MLE is 138/500 = 0.2760. Your gradient descent should land on the same 0.2760, and the NLL drops from 346.57 at p = 0.5 to 294.57 at the MLE. It is not exactly 0.3 because 500 flips are a finite sample.
The Beta(2, 8) prior has its mode at 0.1250. The MAP comes out at 0.2736, between that mode and the MLE but very close to the MLE. With 500 flips the data outvote a prior worth about 8 pseudo-flips.
For the stretch step, the Fisher information at the MLE is about 2502.2, so the standard error is 1/sqrt(2502.2) = 0.0200. The bootstrap standard deviation over 2,000 resamples is 0.0201. A formula and a brute-force resampling agree to the third decimal.
#Quiz
You observe 8 heads in 10 coin flips. What is the MLE for the coin's bias θ?
#Going Further (Optional)
You have the core path now, which takes about 45 minutes. This last section is optional and adds about 6 more: the score function, Fisher information in general, the Cramér-Rao bound and the natural gradient. Skip it on a first read; nothing later depends on it.
#The score function
dℓ/dθ = 0 we solved for the Bernoulli case. In one plain line, E[score] = 0: if the data really come from θ, the score averages to zero. For a Bernoulli, θ × (1/θ) + (1 − θ) × (−1/(1 − θ)) = 1 − 1 = 0. This holds under regularity conditions: the set of possible data values does not depend on θ, and you may differentiate the log-likelihood (and swap the derivative with the sum or integral) as many times as needed.#Fisher information in general
For a Bernoulli flip, E[s²] = θ × (1/θ²) + (1 − θ) × (1/(1 − θ)²) = 1/θ + 1/(1 − θ) = 1/(θ(1 − θ)), and n flips carry n times that, which is the n / (θ(1 − θ)) from the core path. When θ is a vector of many parameters, the squared score becomes an outer product and the information becomes a matrix.
#The Cramér-Rao bound and efficiency
The MLE is asymptotically efficient under the regularity conditions: as n grows, its variance approaches that bound and its distribution approaches a normal curve centered on the true θ. Two cautions. First, "asymptotic" means large n; with little data a shrinkage estimator such as MAP can have a smaller total error than the MLE, because it accepts a little bias for much less variance. Second, the conditions can fail. For data uniform on [0, θ] the possible values depend on θ, and the MLE (the largest observation) does not follow this story.
#The natural gradient
Ordinary gradient descent measures steps as distances in parameter values. For a probabilistic model, what matters is how much the model's predictions change. Two θ values can be far apart in parameter distance but close in distribution, and the reverse. The natural gradient step θ_ = θ_t − η F⁻¹ ∇L multiplies the gradient by the inverse of the Fisher information matrix F, which measures distance between distributions. It is the steepest-descent direction in that distribution-based geometry.
#Connection Across the Math Track
This lesson is the connective tissue:
- Probability & Bayes gave you the formula
p(θ|D) ∝ p(D|θ) p(θ). MAP is what happens when you take the argmax of that posterior. - Random Variables, Expectation & Variance gave you E[X] and Var(X). The Cramér-Rao bound (in Going Further) gives you a lower bound on Var(θ̂) in terms of Fisher information.
- Descriptive Statistics showed you sample variance with the Bessel correction. Now you know WHY that correction exists: MLE for variance is biased; (n−1) makes it unbiased.
- Multivariate Gaussian introduced the Gaussian density and the change-of-variables formula. Gaussian MLE for both μ and Σ pops out of multivariate calculus.
- Information Theory introduced entropy, cross-entropy, and KL divergence. Cross-entropy is exactly negative log-likelihood for categorical models — same number, two names from two communities (information theorists and statisticians).
- Statistical Inference gave you p-values and confidence intervals. For large n and under regularity conditions, the MLE is approximately N(θ_true, 1/I), which gives you a standard error and confidence intervals for the MLE.
- Calculus & Derivatives gave you the gradient and the second derivative. The score is the gradient of the log-likelihood, and Fisher information is its curvature. The natural gradient uses both.
- Matrix Calculus & Backprop gave you the chain rule for vectors. The end-to-end backprop you derived for softmax + cross-entropy is the gradient of the categorical NLL — every modern ML training loop is gradient ascent on the log-likelihood.
Once you see MLE under all the losses, you stop memorizing loss functions. You start deriving them from probabilistic assumptions about your data.