What’s one thing you learned? What’s still confusing?
Vector & Matrix Calculus: Jacobian, Hessian & Backprop Rules
Jacobian, Hessian, key matrix-calculus identities, and a full end-to-end derivation of backprop for a softmax + cross-entropy classifier.
Descriptive Statistics & Distributions
Mean, median, mode, standard deviation, and the distributions that appear everywhere in ML.
Quantiles, Shape & Reading Real Data
Percentiles, quartiles, box plots, skewness, kurtosis, outliers and robust statistics, worked on one real-looking dataset.
Interactive Labs for This Track
Loss Landscape
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Vectors & Matrix Operations
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Probability Distributions
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Ask questions, share insights
loss.backward(). Afterwards, training a network is no longer magic: it is one forward walk and one backward walk over a DAG.x = 2 and y = 3 and compute three things in turn: u = x * y, v = x + y, and L = u * v. The graph has the inputs x and y, then the nodes u, v and L. Going forward: u = 6, v = 5, L = 30. Note that x feeds two nodes (u and v) and so does y.L change if x changes a little? Put the local slope (how much each node's output moves per unit change of one of its inputs) on every arrow:L = u * v gives dL/du = v = 5 and dL/dv = u = 6.u = x * y gives du/dx = y = 3. v = x + y gives dv/dx = 1.x to L. Through u: 5 * 3 = 15. Through v: 6 * 1 = 6. Add the paths: dL/dx = 15 + 6 = 21. For y: through u, 5 * 2 = 10 (since du/dy = x = 2); through v, 6 * 1 = 6; total dL/dy = 16. Check with algebra: L = x²y + xy², so dL/dx = 2xy + y² = 12 + 9 = 21 and dL/dy = x² + 2xy = 4 + 12 = 16. The same numbers.The picture below is this kind of graph. Numbers flow forward along the arrows, and the local slopes sit on the edges. Move an input and watch which node values change, then multiply the slopes along a path to see where each gradient comes from.
What we just did by hand is the general rule: the gradient of the final result with respect to an input equals the sum, over every path from that input to the result, of the product of the local slopes along the path.
loss). The backward pass walks the same graph in reverse. Starting from loss with seed gradient dloss/dloss = 1, each node uses its own local slopes plus its already-computed children's gradients to compute the gradient for its parents. In the example: seed 1 at L, then u gets 5, v gets 6, then x collects 5*3 + 6*1 = 21.1 for a scalar loss) back through the graph. Cost: about one pass total, regardless of how many inputs there are.loss.backward() actually doesloss.backward() walks the computational graph in reverse, applying the chain rule node by node. This lesson's DAG is exactly the graph being walked. Every tensor with requires_grad=True has its .grad field populated with the corresponding partial derivative. The optimizer (SGD, Adam, etc.) then reads .grad and updates the parameter.nn.Embedding(vocab, dim, sparse=True) exists: the gradient of the loss w.r.t. an embedding row is sparse. Only the rows for tokens that appeared in the batch get nonzero gradients (the others were never read during the forward pass, so the chain rule yields zero). A naive dense gradient tensor would be (vocab_size, dim), which for a large vocabulary is a lot of zeros to allocate. The sparse version stores only the rows that changed, cutting memory and compute for large embedding tables.One exercise: a 20-line autograd engine.
loss.backward(). Each Value holds a number, a grad slot, and the nodes it was made from. Each operation records a tiny function _back that pushes gradient to its parents using the local slopes from the worked example. Use += because a node like x can feed several others, and its gradient is the sum over paths. The test graph is the one you solved by hand: x = 2, y = 3, u = x * y, v = x + y, L = u * v.Tests · Verify u = 6, v = 5, L = 30, dL/dx = 21, dL/dy = 16, dL/du = 5, dL/dv = 6, and that the check against the hand calculation prints True.
u, v, L = 6.0 5.0 30.0, then dL/dx = 21.0 dL/dy = 16.0, then dL/du = 5.0 dL/dv = 6.0, then matches hand calculation: True. Those are exactly the numbers from the path sums above. Before you fill in the TODOs the gradients print as 0.0, which is a good sign the skeleton is wired correctly. Notice what reversed(order) guarantees: a node's _back runs only after every node that depends on it has already pushed its gradient in.loss.backward() is just one DAG traversal.Value class with one _back closure per operation, += accumulation and a reversed topological order reproduces loss.backward() on the worked example: dL/dx = 21 and dL/dy = 16.Why is reverse-mode autodiff (backprop) the default for training neural networks instead of forward-mode?