One question, "how surprised should I be?", started the field behind file compression, error correction on noisy links, and the loss function that trains most classifiers. Cross-entropy loss and KL divergence are two of the most used ideas in ML, and both come from one equation Shannon wrote in 1948. Information theory is short, deep, and built on coin flips.
The core path below takes about 30 minutes: surprise, entropy, cross-entropy, KL divergence, why KL is never negative, and why the usual classifier loss is the negative log-likelihood of the true class. The final section, Going Further (Optional), is extra reading on top of that (Jensen's inequality, mutual information, the ELBO, channel capacity and more). Skip it on a first pass.
Learning Objectives
After this lesson, you will be able to:
Measure how surprising an event is with a simple formula: rare events carry more information, common events carry less
Calculate entropy, the average surprise of a random process, and see why a fair coin is maximally unpredictable while a rigged one is not
Compute cross-entropy, the average surprise when you hold the wrong beliefs, and see why it is the standard error formula for classifiers
Understand KL divergence as the extra surprise from using the wrong distribution, check by hand that it is never negative, and explain why minimizing cross-entropy and minimizing KL divergence are the same job
Show that cross-entropy loss is the negative log-likelihood of the true class
Probability told you how likely things are. This lesson asks how much you LEARN when something actually happens. It is the bridge between uncertainty and knowledge, and it is where the standard classifier loss comes from.
Claude Shannon, working at Bell Labs in 1948, asked a deceptively simple question: "How do you measure information?" His answer created an entire field and, as it turns out, supplied the error formula used by most modern classifiers.
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type the small calculations in this lesson yourself.
Start with three events and count the surprise in coin flips.
A fair coin lands heads. Probability 1/2. One flip, 1 bit of surprise by definition.
Three fair coins all land heads. Probability 1/2 × 1/2 × 1/2 = 1/8. Three flips' worth: 3 bits.
A fair die shows a six. Probability 1/6. That is a bit more surprising than 1/4 (2 bits) and less than 1/8 (3 bits): it comes out at 2.585 bits.
The pattern: probability 1/2 gives 1 bit, 1/4 gives 2, 1/8 gives 3. Each halving of the probability adds one bit. The function that turns "1/8" into "3" is the base-2 logarithm with a minus sign, because log₂(1/8) = -3. That is the definition of information content (also called "surprise" or "self-information"):
I(x)=−log2P(x)
Why the logarithm? Two reasons:
Additivity: Learning two independent facts should give you the sum of each fact's information. Since P(A and B) = P(A) × P(B) for independent events, the log turns multiplication into addition: I(A and B) = I(A) + I(B). The three coins above are exactly this.
Units: The base of the log sets the unit. Base 2 gives bits, the natural log gives nats. One bit is the information in one fair coin flip.
What Do You Think?
Two towns forecast 'sunny tomorrow'. In Dryville it is sunny on 82 of every 100 days. In Cloudton it is sunny on 41 of every 100 days. Which forecast carries more information?
The Cloudton forecast carries more information. A sunny day there is less expected, so learning about it tells you more: 1.29 bits against 0.29 bits.
Quick check
A 4-class classifier outputs probabilities [0.97, 0.01, 0.01, 0.01]. The true label turns out to be class 1. What is the information content (in bits) of this observation under the model's distribution?
Surprise belongs to one outcome. A random process has many outcomes, so ask the average question: on average, how surprised will I be by the next result? Weight each outcome's surprise by how often it happens.
Fair coin. Both outcomes have probability 0.5 and surprise 1 bit. Average: 1 bit.
Biased coin, 90% heads. Heads is expected, so its surprise is only -log₂(0.9) = 0.152 bits. Tails is rare, so its surprise is -log₂(0.1) = 3.322 bits. Heads happens 90% of the time and tails 10%, so the average is
Suppose a 4-class image classifier outputs the following probabilities for one image: $P = [0.7, 0.2, 0.07, 0.03]$.
H(P)=−(0.7log20.7+0.2log2
The maximum possible entropy for 4 classes is $\log_2 4 = 2.0$ bits (the uniform distribution). This classifier's 1.245 bits is below that cap because class 0 dominates.
Maximum entropy occurs when all outcomes are equally likely. A fair coin (1 bit) has more entropy than a biased coin (0.469 bits). A fair 6-sided die (log₂ 6 = 2.585 bits) has more entropy than a loaded die.
Minimum entropy (zero) occurs when one outcome has probability 1. Zero uncertainty: you know exactly what will happen.
Entropy measures uncertainty, not randomness. High entropy means the next outcome is hard to predict.
Now see it. In the picture below, drag the probabilities and watch the surprise bars: the total area of the bars is the entropy, and a marker shows the largest entropy possible for that many outcomes. Try to arrange 4 outcomes so the entropy reaches the marker.
Entropy assumed you know the true probabilities. Now suppose you do not. A coin is fair (the truth, P = [0.5, 0.5]) but your model believes Q = [0.9, 0.1]: it thinks heads comes up 90% of the time.
How surprised is the model, on average, by what really happens? Reality gives heads half the time and tails half the time, but the surprise is scored with the model's beliefs:
Heads happens half the time, and the model's surprise is -log₂(0.9) = 0.152 bits.
Tails happens half the time, and the model's surprise is -log₂(0.1) = 3.322 bits.
Average: 0.5 × 0.152 + 0.5 × 3.322 = 1.737 bits. A model with the right beliefs would average 1 bit (the fair coin's entropy). The wrong beliefs cost 0.737 extra bits per flip.
That average is the cross-entropy: the average surprise when events come from the true distribution P but you score them with your model's distribution Q.
H(P,Q)=−x∑P(x)log2Q(x)
What it measures, precisely: cross-entropy is the average surprise under the wrong beliefs Q. Entropy H(P) is the part of that surprise that is unavoidable even with perfect beliefs. The extra on top is the KL divergence, defined in the next section.
Quick check
The true distribution is P = [0.5, 0.5]. Your model predicts Q = [0.9, 0.1]. Is the cross-entropy H(P, Q) larger or smaller than H(Q, P)?
For a single training example with true class k, the true distribution P is a one-hot vector: all zeros except a 1 at position k. Every term with P = 0 vanishes, and cross-entropy collapses to one term:
L=−log2Q(ytrue)
This is also the negative log-likelihood of the true class. The likelihood of the data is the probability the model assigned to what actually happened. Take three training examples where the model gave the true class probabilities 0.8, 0.5 and 0.1:
Likelihood of all three together: 0.8 × 0.5 × 0.1 = 0.04.
Negative log of it: -log₂(0.04) = 4.644 bits.
Because a log turns the product into a sum, that is -log₂ 0.8 - log₂ 0.5 - log₂ 0.1 = 0.322 + 1 + 3.322, the sum of the three per-example losses. Divide by 3 and the average loss per example is 1.548 bits.
So averaging the cross-entropy loss over a dataset and maximizing the likelihood of the labels are the same job seen from two sides. Minimizing one maximizes the other.
Perplexity. Language models are scored with perplexity: 2 raised to the average cross-entropy in bits, or equivalently e raised to the average cross-entropy in nats. Frameworks report the loss in nats, so check the unit first. An average of 2 bits per token gives a perplexity of 2² = 4, as uncertain as choosing between 4 equally likely tokens. The same model reports a loss of 1.386 nats, and e^1.386 = 4 again.
Back to the coin. You computed three numbers: the entropy of the truth H(P) = 1 bit, the cross-entropy H(P, Q) = 1.737 bits, and their difference, 0.737 bits. That difference is the extra surprise from believing Q when reality is P. It has a name, the KL divergence (Kullback-Leibler divergence):
DKL(P∥Q)=x∑P(x)log2Q(x)P(x)
What Do You Think?
Suppose your model's predictions Q exactly match the true distribution P. What is KL(P‖Q)?
The relationship that matters most: Cross-Entropy = Entropy + KL Divergence.
H(P,Q)=H(P)+DKL(P∥Q)
For the coin: 1.737 = 1 + 0.737. Entropy H(P) depends only on the truth, not on your model, so minimizing cross-entropy is exactly minimizing KL divergence. Training pushes the model's distribution as close to the truth as possible.
Drag the bars of Q below and read H(P), H(P, Q) and the KL. H(P) never moves, because it does not depend on Q. Watch what happens to the other two when Q reaches P.
Loading visualization...
The "wrong beliefs" tab of the surprise-bars picture above shows the same three numbers.
Could a wrong model ever do better than the truth, giving a negative KL? No. Here is a short check using one fact about the natural logarithm: for every x > 0, ln x ≤ x − 1. (The straight line y = x − 1 touches the curve y = ln x at x = 1 and lies above it everywhere else. Check: ln 0.5 = -0.693 ≤ -0.5, and ln 2 = 0.693 ≤ 1.)
Work in natural logs for the proof (that only rescales KL by the positive number ln 2). Apply the fact with x = Q/P to each outcome where P > 0:
-KL = Σ P · ln(Q/P) ≤ Σ P · (Q/P − 1) = Σ (Q − P).
Σ P = 1, and Σ Q over those outcomes is at most 1. So Σ (Q − P) ≤ 0.
Therefore KL ≥ 0, and it equals 0 only when Q = P.
Check with the coin: -KL = 0.5 · ln(0.9/0.5) + 0.5 · ln(0.1/0.5) = 0.5 × 0.588 + 0.5 × (−1.609) = −0.511 nats (that is 0.737 bits × 0.693), and the bound says it must be at most 0. It is.
The support condition. The argument needs Q to be positive wherever P is positive. If the truth sometimes produces an outcome that your model calls impossible (Q = 0 while P > 0), the term P · log(P/0) is infinite: KL is infinite. A model that says "never" about something that happens is infinitely surprised. Outcomes with P = 0 contribute nothing (by convention 0 · log 0 = 0), whatever Q says there.
Since KL ≥ 0, cross-entropy is at least the entropy: H(P, Q) ≥ H(P). The smallest loss any classifier can reach is the entropy of the true labels, not zero (unless every label is certain). With one-hot labels, H(P) = 0 and the KL divergence is the loss itself, -log Q(y_true).
KL is not symmetric. For the coin, D_KL(P‖Q) = 0.737 bits but D_KL(Q‖P) = 0.9 · log₂(0.9/0.5) + 0.1 · log₂(0.1/0.5) = 0.531 bits. Swapping the arguments changes the answer, so KL is a measure of divergence but not a distance. Going Further explains what the two directions each reward.
Everything in the core path fits in a few lines of Python. You will compute entropy, cross-entropy and KL for three models of one true distribution, then the negative log-likelihood and perplexity of a small dataset, and check that the identities hold.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
Tests · Verify H(P) = 1.1568; model A cross-entropy 1.1896 and KL 0.0328; model B cross-entropy 2.7219 and KL 1.5651; model C KL = 0 and cross-entropy = entropy; cross-entropy = entropy + KL for all three models; likelihood 0.04, -log2 = 4.6439, average 1.5480 equal to the mean of per-example losses; perplexity 2.9240; a zero in Q where P > 0 gives infinity.
Running the solution prints H(P) = 1.1568 bits. Model A (pretty good) has cross-entropy 1.1896 and KL 0.0328. Model B (confident and wrong) has cross-entropy 2.7219 and KL 1.5651. Model C (perfect) has cross-entropy 1.1568, equal to the entropy, and KL exactly 0. For the three-example dataset the likelihood is 0.04, the negative log of it is 4.6439 bits, the average loss is 1.5480 bits per example, and the perplexity is 2.9240.
#Live: Entropy, Cross-Entropy, and KL Side by Side
Run this cell to compute H(p), H(p, q), and KL(p‖q) for two distributions over five classes, plot them side by side, then verify that KL is NOT symmetric.
Rare events carry more information. Information content is -log(P), so unlikely events are highly informative while certain events tell you nothing new
Entropy measures average surprise. A uniform distribution has maximum entropy (hardest to predict), while a distribution concentrated on one outcome has minimum entropy (easiest to predict)
Cross-entropy is the average surprise under your model's beliefs while events come from the truth. For a classifier it equals the negative log-likelihood of the true class, so minimizing it maximizes the likelihood of the labels
Confident wrong predictions are punished heavily, since the loss is -log(p) for the true class. That does not guarantee calibration: networks trained with cross-entropy are often overconfident
Cross-entropy equals entropy plus KL divergence, and KL is never negative (given Q is positive wherever P is). So the loss can never beat the entropy of the truth, and minimizing cross-entropy is minimizing KL
A language model assigns probabilities: P('the')=0.4, P('a')=0.3, P('cat')=0.2, P('dog')=0.1. The true next word is 'cat'. What is the cross-entropy loss?
Everything above is the core path. What follows is extra and entirely optional, roughly another 15 minutes if you read it all. Nothing in the core path depends on it, but Jensen's inequality comes back in the next lesson, and the first item here is where it is taught. Each item is folded away, so open only what you want.
A Mathematical Theory of Communication
Claude Shannon (1948)
The paper that started it all. Surprisingly readable. Introduces entropy, channel capacity, and the fundamental limits of communication. Every ML practitioner should read at least the first few sections.
⚡ Playground:Probability Explorer → — change the probability distribution and watch entropy rise and fall.
Next up: Probability Inequalities & Monte Carlo. Entropy and KL divergence told you how surprising data are. Next you will bound how likely an average is to land far from its true value, and estimate expectations by sampling. Jensen's inequality, taught in this lesson's Going Further section, is used there rather than taught again.