Random Variables, Expectation & Variance
After this lesson, you will be able to:
- Understand random variables as rules that turn random outcomes into numbers, and tell apart the two types: countable (like dice) and continuous (like temperature)
- Calculate the expected value (long-run average) of a random process and understand why it matters even when no single outcome equals the average
- Measure how spread out a random process is using variance and standard deviation, and see how AI models use these same ideas in their error-measuring formulas
- Apply linearity of expectation to compute E[aX + bY] without needing joint distributions, and explain why Var(X + Y) is not the same formula when X and Y are dependent
Before You Start
#From Outcomes to Numbers
Random does not mean unpredictable. A single coin flip is random, but flip a coin 10,000 times and you can predict almost exactly how many heads you will get. This lesson teaches you how to turn randomness into reliable predictions, which is exactly what every AI model does.
There are two flavors:
- Discrete random variables take countable values: die rolls (1-6), number of spam emails per day, pixel values (0-255)
- Continuous random variables take any value in a range: temperature, height, neural network weights, loss values
#Discrete: The PMF
For a fair die:
| x | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| P(X = x) | 1/6 | 1/6 | 1/6 | 1/6 | 1/6 | 1/6 |
Rules: every probability is between 0 and 1, and they all sum to 1.
For a loaded die that favors 6:
| x | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| P(X = x) | 1/10 | 1/10 | 1/10 | 1/10 | 1/10 | 1/2 |
A PMF answers questions by adding table entries. For the loaded die, the chance of rolling at least a 5 is P(X = 5) + P(X = 6) = 0.1 + 0.5 = 0.6. Check the sum rule too: 0.1 + 0.1 + 0.1 + 0.1 + 0.1 + 0.5 = 1.
The PMF completely describes the discrete random variable. If you know the PMF, you know everything.
#Try the distribution explorer
#Continuous: The PDF
A continuous random variable, like a temperature or a waiting time, can land anywhere in a range, so it has infinitely many possible values. That breaks the table idea. The chance of hitting one exact value, such as exactly 21.3000... degrees, is zero. What carries probability is a range of values.
Worked example. Suppose X is equally likely to be anywhere between 0 and 0.5, which makes its curve a flat rectangle. The base is 0.5 and the total area must be 1, so the height must be 2. The chance that X lands between 0.1 and 0.3 is the area of the slice of rectangle with base 0.2 and height 2: 0.2 times 2 = 0.4.
Mathematicians write "the exact area under the curve between a and b" with an integral sign. It is the same area as the histogram bars, in the limit of infinitely thin bars:
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
#The CDF: Cumulative Distribution Function
P(X <= x) answers: "What is the probability that X is at most x?" It is a running total of probability, added up from the far left to x.Worked example with the fair die. F(4) = P(X is 1, 2, 3 or 4) = 4/6, and F(2) = 2/6. The chance of rolling a 3, 4 or 5 is F(5) minus F(2) = 5/6 - 2/6 = 3/6 = 0.5. One subtraction, no list of cases.
For discrete variables, the CDF is a staircase function that jumps at each possible value. For continuous variables, it is a smooth S-shaped curve that goes from 0 to 1.
The CDF is incredibly useful because:
- P(X > x) = 1 - F(x)
P(a < X <= b)= F(b) - F(a)- The median is the value where F(x) = 0.5
CDFs come back when you compute p-values, which a later lesson covers.
#Expectation: The Long-Run Average
A fair six-sided die has E[X] = 3.5. Can you ever actually roll 3.5?
Do it with numbers first. For the fair die, multiply each face x by its probability, and add up the results. The last column squares the face before weighting; the variance section below uses it.
| x | P(X = x) | x times P | x squared times P |
|---|---|---|---|
| 1 | 1/6 | 0.1667 | 0.1667 |
| 2 | 1/6 | 0.3333 | 0.6667 |
| 3 | 1/6 | 0.5 | 1.5 |
| 4 | 1/6 | 0.6667 | 2.6667 |
| 5 | 1/6 | 0.8333 | 4.1667 |
| 6 | 1/6 | 1.0 | 6.0 |
| Total | 1 | 3.5 | 15.1667 |
You can never roll 3.5, because it is not even on the die. But if you roll 1,000 times and average the results, you will get very close to 3.5. The expected value is about the long run, not any single trial.
Now the loaded die, with the same columns. Watch the numbers move.
| x | P(X = x) | x times P | x squared times P |
|---|---|---|---|
| 1 | 0.1 | 0.1 | 0.1 |
| 2 | 0.1 | 0.2 | 0.4 |
| 3 | 0.1 | 0.3 | 0.9 |
| 4 | 0.1 | 0.4 | 1.6 |
| 5 | 0.1 | 0.5 | 2.5 |
| 6 | 0.5 | 3.0 | 18.0 |
| Total | 1 | 4.5 | 23.5 |
What the table did, written as a formula, is "value times its probability, summed over every value":
Drag the fulcrum along the beam until it balances, and read off where it rests: that is E[X]; then flip to the squares tab to see variance drawn as squares, which the next section computes.
#Linearity of Expectation
This is the single most useful property in probability. It says:
This works even when X and Y are dependent. It is one of the rare properties in probability that does not require independence. This is why it is so powerful.
#Variance: How Spread Out Are the Outcomes?
Do it with numbers first, on the fair die. Its center is 3.5. The distance from 3.5 to each face is -2.5, -1.5, -0.5, 0.5, 1.5, 2.5. Squaring those distances (so that left and right do not cancel) gives 6.25, 2.25, 0.25, 0.25, 2.25, 6.25. Their average is 17.5/6 = 2.9167. That average squared distance from the mean is the variance.
There is a shortcut that uses the last column of the tables above. The average squared face is E[X^2] = 15.1667, and the squared average is 3.5 squared = 12.25. The difference is 15.1667 - 12.25 = 2.9167, the same answer with no distances. For the loaded die: 23.5 - 4.5 squared = 23.5 - 20.25 = 3.25. The loaded die is both higher and a little more spread out, because the 6 sits further from the center.
Written as a formula:
#Variance Properties
- Var(aX + b) = a^2 Var(X). Adding a constant does not change spread, but scaling by a multiplies variance by a^2.
Adding two random variables is trickier, because the spread of X + Y depends on whether X and Y move together. We need one new quantity.
Now the rule for sums:
- Var(X + Y) = Var(X) + Var(Y) + 2 Cov(X, Y) holds in general.
- Var(X + Y) = Var(X) + Var(Y) when X and Y are uncorrelated (Cov = 0).
Check it on the example. Var(X) is the average of the squared deviations (1, 0, 1), which is 2/3, and Var(Y) is 2/3 too. The sums X + Y are 3, 3 and 6, with average 4, so their deviations are -1, -1 and 2, their squares 1, 1 and 4, and Var(X + Y) = 6/3 = 2. The formula gives 2/3 + 2/3 + 2 times 1/3 = 2. It matches. Leaving out the covariance term would give 4/3, which is wrong.
Standardizing covariance gives the correlation coefficient, which a later lesson, Correlation, Causation & Simpson's Paradox, develops.
Independent and uncorrelated sound like the same thing. They are not, and a three-point example shows why. Let X be -1, 0 or 1, each with probability 1/3, and let Y = X squared.
- Y is 1, 0, 1 for those three cases, so E[Y] = 2/3. Also E[X] = 0.
- The products XY are -1, 0 and 1, so E[XY] = 0.
- Cov(X, Y) = E[XY] - E[X] E[Y] = 0 - 0 times 2/3 = 0. They are uncorrelated.
- Yet Y depends on X completely: once you know X, you know Y exactly. Knowing X = 0 tells you Y = 0, and any other X tells you Y = 1.
Covariance only detects straight-line tendencies, and Y = X squared is a U shape. The continuous version, X ~ Uniform(-1, 1) (meaning X is equally likely anywhere between -1 and 1) with Y = X squared, gives the same zero by symmetry.
A random variable X has Var(X) = 4. What is Var(3X − 7)?
#Compute E[X], Var[X], and watch the Law of Large Numbers
Theory says E[X] = 3.5 for a fair die. The Law of Large Numbers says that the more times you repeat a random experiment, the closer the average of your results gets to E[X]. The cell below shows it experimentally, and it lets you flip to a loaded die and see exactly how the long-run mean shifts.
#Interactive: Explore Distributions
The explorer below draws four continuous distributions and shows each one's mean and variance.
The binomial is not in the explorer above, so here is the one from the PMF section again, which also has a Mean readout. Set n = 20 and p = 0.3 and check that E[X] = n times p = 6.
#Connection to ML: Loss Functions ARE Expectations
#Try It Yourself
You have computed the die numbers by hand. Now let code do the same arithmetic, then check it three ways: against a simulation, against the linearity and scaling rules, and against the covariance example. Everything you need is in the lesson above.
Tests · Verify E[X] = 3.5 and Var = 2.9167 for the fair die, E[X] = 4.5 and Var = 3.25 for the loaded die, simulated means within 0.05 of theory, E[2X+3] = 12 and Var(3X-7) = 29.25 for the loaded die, Cov = 1/3 with Var(X+Y) = 2, and Cov(X, X^2) = 0.
Running the solution prints these values. The fair die gives E[X] = 3.5000, Var = 2.9167 and Std = 1.7078, with a simulated mean of 3.4973 and variance 2.9143. The loaded die gives E[X] = 4.5000, Var = 3.2500 and Std = 1.8028, with a simulated mean of 4.4965 and variance 3.2409. For the loaded die E[2X + 3] = 12.0, equal to 2 times 4.5 + 3, and Var(3X - 7) = 29.25, equal to 9 times 3.25. The covariance example prints Cov(X, Y) = 0.3333, and Var(X + Y) = 2.0 matches the formula, while leaving out the covariance term gives 1.3333. The stretch case prints Cov(X, X^2) = 0.0 even though Y is a function of X.
#Key Takeaways
- Random variables map outcomes to numbers — they transform real-world randomness into mathematical objects that ML can work with; discrete variables have PMFs, continuous variables have PDFs
- E[X] is the long-run average — you may never observe E[X] in a single trial (like rolling 3.5 on a die), but over many trials the average converges to it
- Linearity of expectation always holds — E[aX + bY] = aE[X] + bE[Y] regardless of independence, making it the most powerful tool in probability
- Variance measures spread — Var(X) = E[X^2] - (E[X])^2 quantifies how far outcomes deviate from the mean; unlike expectation, variance is NOT linear unless variables are independent
- ML loss functions are expectations — training minimizes E[L(y, f(x))], and batch normalization computes E[X] and Var(X) over mini-batches
#Quick Check
A fair die has E[X] = 3.5 and Var(X) = 2.917. If you define Y = 2X + 1, what is E[Y]?
You now have the toolbox for the checkpoint: counting and probability, Bayes' theorem, random variables, expectation, variance and covariance. Next you will use all of it together on a short review.