What’s one thing you learned? What’s still confusing?
Vectors & Spaces
Understand vectors as geometric objects with direction and magnitude, and learn how vector spaces form the foundation of ML data representation.
Matrices as Transformations
See matrices as space-warping machines. Every neural network layer is a matrix multiplication — learn why that matters.
Eigenvalues & SVD
Discover the special directions that matrices stretch without rotating, and how SVD decomposes any transformation into simple pieces.
Interactive Labs for This Track
Loss Landscape
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Vectors & Matrix Operations
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Probability Distributions
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Ask questions, share insights
{1, 2, 3} and B = {3, 4}.| Symbol | Read as | Example with the answer |
|---|---|---|
| ∈ | "is in" (and ∉ means "is not in") | 2 ∈ A is true. 5 ∈ A is false, so 5 ∉ A. |
| ∪ | union: everything in either set | A ∪ B = {1, 2, 3, 4} |
| ∩ | intersection: only what is in both | A ∩ B = {3} |
| ∀ | "for all" | "For all x in A, x > 0" is true, because 1, 2 and 3 are all above 0. |
| ∃ | "there exists" | "There exists x in A with x > 2" is true, because x = 3 works. |
| such that | written as a colon or a bar | {x ∈ A : x > 1} reads "the x in A such that x is above 1", which is {2, 3}. |
When a later lesson says an event is a set of outcomes, "A or B" will be the union and "A and B" will be the intersection. Python has the same ideas built in, so you can check each row.
Start with a small sum worked by hand: add the first four square numbers, 1 squared, 2 squared, 3 squared and 4 squared. Keep a running total as you go.
| i | Term (i squared) | Running total |
|---|---|---|
| 1 | 1 | 1 |
| 2 | 4 | 1 + 4 = 5 |
| 3 | 9 | 5 + 9 = 14 |
| 4 | 16 | 14 + 16 = 30 |
The answer is 30. Sigma is the short way to write that whole table. Here is the plainest case, adding four numbers x_1, x_2, x_3 and x_4:
Here is the same sum of i squared as a plain Python loop. The total starts at 0, and each pass through the loop adds the next term:
total = 0
for i in range(1, 5): # range(1, 5) counts 1, 2, 3, 4 and stops before 5
total = total + i ** 2
print(total) # 30sum(...) adds up everything it is given, so sum(i ** 2 for i in range(1, 5)) is the one-line version of the loop. Later you will also see zip(a, b), which walks along two lists at once and hands you the pair in each position: zip([1, 2, 3], [4, 5, 6]) gives (1, 4), (2, 5), (3, 6).The stepper below does the same table for several expressions.
Pick an expression, press the step button, and watch each term join the running total, then switch to a pi expression and see the running product.
Each of these is a sum in disguise. Work through them once and the later lessons will read easily.
The capital Greek letter pi, written with its big product symbol, means "multiply these" in exactly the same way sigma means "add these". The index, bounds and subscripts work identically.
Now the bridge from Part 1. Take the log of a pi expression and the product rule from the logs section turns it into a sigma expression:
That one line is the whole reason maximum likelihood methods are written as sums of log-probabilities.
A formula is a sentence. It has a handful of symbols, and once you know the alphabet you can read it left to right. Here is the alphabet that ML uses most.
| Symbol | Name | Typical meaning |
|---|---|---|
| α (alpha) | alpha | a learning rate or a small tuning constant |
| β (beta) | beta | a coefficient, or a momentum setting |
| θ (theta) | theta | the parameters (weights) of a model |
| μ (mu) | mu | the mean of a distribution |
| σ (sigma, lowercase) | sigma | a standard deviation. Do not confuse with capital Σ, which means "sum" |
| λ (lambda) | lambda | a regularisation strength, or a rate |
| ε (epsilon) | epsilon | a tiny number, or noise |
| x with a hat, x̂ | x-hat | an estimate or prediction of x |
| x with a bar, x̄ |
Two reading habits help. First, the same letter can mean different things in different books, so check the line near the formula that says what each symbol stands for. Second, brackets and bars group things. Work the inside first.
Softmax is the function that turns a list of raw scores z into probabilities. Before the formula, build it by hand for just two scores, z = [1, 0].
The two answers add to 1, and the larger score got the larger share. The formula is that recipe written for any number K of scores (K is just how many classes there are).
Read it word by word:
| Piece | What it computes | Value at x = 1 |
|---|---|---|
| x minus mu | distance from the centre | 1 |
| squared | so left and right count the same | 1 |
| divided by two sigma squared | scale by the width | 1 / 2 = 0.5 |
| minus, then exp | 1 at the centre, shrinking outward | exp(-0.5) = 0.6065 |
| times one over sigma root two pi | a constant (pi is the circle number, 3.14159) | 0.3989 x 0.6065 = 0.2420 |
Stack those five pieces together and you have the formula.
Word by word, working from the middle outward:
| Class | Score z | exp(z) | Probability |
|---|---|---|---|
| A | 2.0 | 7.389 | 7.389 / 11.213 = 0.659 |
| B | 1.0 | 2.718 | 2.718 / 11.213 = 0.242 |
| C | 0.1 | 1.105 | 1.105 / 11.213 = 0.099 |
| Total | 11.213 | 1.000 |
The biggest score gets the biggest share, but every class keeps some probability. Because exp grows quickly, a score gap of one point already gives A nearly three times B's share.
Change one score in the widget above and watch all the probabilities move together, then add 100 to every score and see that nothing changes.
That add-100 switch also exposes a practical trap. The probabilities do not change, but the naive exp of a large score can overflow. A 32-bit float, common on GPUs, tops out near e^88.7, so exp(102) is already inf there, and inf divided by inf is NaN. Plain Python floats are 64-bit and reach much further: math.exp(709) is about 8.2e307, and math.exp(710) raises an OverflowError (NumPy returns inf instead). So the widget's overflow demo is a 32-bit story, and in Python the same trouble starts only near exp(710). The standard fix works for both. Subtract the largest score from every score before applying exp. The largest exponent becomes 0, so nothing overflows, and the probabilities are unchanged because exp(z - m) = exp(z) / exp(m) and the common factor cancels in the division.
Notice the same σ letter again. Here, sigma(t) is the sigmoid function. Context tells you which sigma you are looking at: capital Σ is a sum, lowercase σ is usually a standard deviation, and σ followed by a bracket is often the sigmoid.
Why ML cares: softmax and sigmoid sit at the end of nearly every classifier, and cross-entropy then takes the log of the probability they produce. Exp and log are two halves of one pipeline.
For a straight line like f(x) = 2x + 1 the slope is the same everywhere: 2. For a curve it changes. Take f(x) = x squared and look at the point x = 3 with a shrinking run:
| Between | Rise | Run | Slope |
|---|---|---|---|
| x = 3 and x = 4 | 16 - 9 = 7 | 1 | 7.0 |
| x = 3 and x = 3.1 | 9.61 - 9 = 0.61 | 0.1 | 6.1 |
| x = 3 and x = 3.01 | 9.0601 - 9 = 0.0601 | 0.01 | 6.01 |
As the run shrinks, the slope settles toward 6. That settled value is the slope of the curve exactly at x = 3. Chasing it as the run shrinks to nothing is the idea behind derivatives, and it is where the next lessons pick up. In ML, the slope of the loss tells a model which way to nudge its weights to improve.
Tests · Verify the naive product equals 0.0 (underflow), the log-likelihood is a normal finite negative number, the average per event is the log-likelihood divided by 400, and the cross-entropy for p = 0.7 on the true class is about 0.3567.
When you run the solution you should see the naive product print as 0.0, so the "product is zero" check prints True. The log-likelihood is -794.38, the average per event is -1.9859, and the cross-entropy is 0.3567.
What does the sum of i squared, for i running from 1 to 4, equal?
| x-bar |
| the mean of x |
| x_i | subscript | the i-th item, or a label such as x_train |
| x squared, x to the n | superscript | a power, or sometimes an index in brackets such as x(i) |
| a over b | fraction bar | divide the top by the bottom |
| arg max | argument of the maximum | the input that gives the largest output, not the largest output itself |