Vector & Matrix Calculus: Jacobian, Hessian & Backprop Rules
.backward() is doing exactly what's in this lesson, just with bigger matrices.After this lesson, you will be able to:
- Compute the Jacobian of a vector-valued function and apply the vector-valued chain rule J_h = J_g · J_f to chain transformations
- Read a Hessian as a table of curvatures and say what the signs of its eigenvalues mean
- Apply the load-bearing matrix-calculus identities (∂(Ax)/∂x = A, ∇(x^T A x) = (A+A^T)x, ∇tr(AB) = B^T) with one convention: a gradient always has the shape of the variable
- Explain why backprop starts at the loss and runs backward through the graph
- Derive the softmax + cross-entropy gradient ∂L/∂z = p − y_one_hot from scratch, then use it in a training loop whose hand-written gradient you check against finite differences
Before You Start
#From Scalars to Vectors to Matrices
Backprop is not magic. It is matrix calculus, applied recursively, with a clever trick. The previous lesson taught you what a derivative is. This one teaches you what a derivative of a vector function with respect to a vector is, and why the chain rule, generalised properly, is exactly what makes neural networks trainable.
f'(x) is a number. The gradient ∇f(x) (when f returns one number) is a vector. The Jacobian (when f returns a whole vector) is a table of partial derivatives. The Hessian (second derivatives of a one-number function) is an n × n table. As you generalise from one input to many, and one output to many, derivatives become matrices, but the chain rule survives, dressed up as matrix multiplication.Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
#The Jacobian: Many In, Many Out
f that takes a vector of n numbers and returns a vector of m numbers. Its Jacobian J is the table whose entry in row i, column j is the partial derivative of the i-th output with respect to the j-th input:#Worked Example: A 2-In, 2-Out Function
f(x, y) = (x²y, x + y). We want the Jacobian at the point (x, y) = (2, 3).The four partials:
∂f₁/∂x = ∂(x²y)/∂x = 2xy. At (2, 3):2·2·3 = 12.∂f₁/∂y = ∂(x²y)/∂y = x². At (2, 3):4.∂f₂/∂x = ∂(x+y)/∂x = 1.∂f₂/∂y = ∂(x+y)/∂y = 1.
So the Jacobian at (2, 3) is:
The rows of the Jacobian are gradients of individual outputs. The columns describe how each input affects the entire output vector. Both readings matter, and different parts of backprop use different ones.
Two shapes to keep apart
A Jacobian is a table with one row per output and one column per input. A gradient is a different object: it belongs to a function that returns one number, and it has the same shape as the variable you differentiate with respect to. The gradient with respect to a column of n numbers is a column of n numbers. The gradient with respect to a 3 × 4 weight matrix is a 3 × 4 matrix.
param -= lr * grad just works. A function with one output has a Jacobian that is a single row, and its gradient is that row stood up as a column. Some books write gradients as rows instead. If a formula in a paper looks off by a transpose, that is usually the reason.Switch between vector fields and scalar fields to see the Jacobian (a table of derivatives of a vector field) and the gradient plus Hessian (derivatives of a scalar field) computed live at any point you click.
A function f: R³ → R² takes a 3-vector and returns a 2-vector. What shape is its Jacobian J_f?
For the linear map y = W x with W ∈ R^{m × n}, what is the Jacobian ∂y/∂x?
#The Vector-Valued Chain Rule
h(x) = g(f(x)), the Jacobian of the composition is the product of Jacobians:#Worked Example: Composing Two Functions
f(x, y) = (x + y, xy) and g(u, v) = u² + v. Then h = g ∘ f is a function R² → R.J_f = [[1, 1], [y, x]]. (At (2, 3): [[1, 1], [3, 2]].)J_g = [2u, 1]. (One row, because g has one output.)J_g = [10, 1].J_h = J_g · J_f = [10·1 + 1·3, 10·1 + 1·2] = [13, 12].h(x, y) = (x+y)² + xy, so ∂h/∂x = 2(x+y) + y and ∂h/∂y = 2(x+y) + x. At (2, 3): [2·5 + 3, 2·5 + 2] = [13, 12]. ✓This is the entire mathematical content of backpropagation. A neural network is a long composition of functions, and the gradient of the loss with respect to any parameter is a long chain of Jacobian matrix products.
∇ₓL = Jᵀ · ∇ᵧL. Going backward means multiplying by the transpose of each layer's Jacobian.Watch the chain rule fire through a stack of layers. Input flows forward through three function blocks, then the gradient flows back as a product of Jacobians at each block.
Now animate the full forward-then-backward sweep. The animation makes it clear why people call it "backprop": the loss signal genuinely walks right-to-left across the graph.
#The Hessian: Curvature in n Dimensions
f that returns one number, the first derivatives form the gradient. The second derivatives form the Hessian H, an n × n table where entry (i, j) is how the i-th slope changes when you move in direction j:f(x, y) = x² + 3xy + y³ at the point (1, 2). Differentiate twice: ∂²f/∂x² = 2, ∂²f/∂x∂y = 3, ∂²f/∂y² = 6y = 12. SoThe eigenvalues of H (from the Eigenvalues & SVD lesson) say how the surface bends along H's special directions. Here they are about 1.17 and 12.83, both positive, so the surface curves up in every direction at this point, like a bowl. A negative eigenvalue means it curves down along that direction. A mix of positive and negative is a saddle: up in some directions, down in others.
At a point where the Hessian has eigenvalues {+3, +1, −0.2}, what does the surface look like?
#The Identity Toolkit
These are the matrix-calculus identities you will use weekly. They all follow the shape convention from the Jacobian section: the gradient of a one-number function has the same shape as the variable.
Ax returns a vector:The rest are gradients of a single number, so each result has the shape of the variable. The simplest is an inner product:
Before the next identity, make a guess.
Let A = [[1, 2], [3, 4]] and x = [1, 1]. The number x^T A x equals 10. What is the gradient of x^T A x with respect to x at this point?
Do not take these on trust. The cell below nudges each entry of the variable up and down and compares the slope with the formula, using a deliberately non-symmetric A.
(3,) for the vector cases and (3, 3) for the trace case: each gradient has the shape of its variable.#The Outer Product That Backprop Loves
y = Wx with W = [[1, 0, 2], [−1, 1, 0]] (2 × 3) and x = [1, 2, 3], so y = [7, 1]. Suppose the loss is L = 0.5·y₁ − 1.0·y₂, so the upstream gradient is ∂L/∂y = [0.5, −1.0].W_ij only affects output y_i, and it does so by multiplying x_j. So the gradient entry for W_ij is (upstream gradient i) times (input j):∂L/∂W = [[0.5·1, 0.5·2, 0.5·3], [−1·1, −1·2, −1·3]] = [[0.5, 1, 1.5], [−1, −2, −3]].∂L/∂x = Wᵀ · [0.5, −1] = [1.5, −1, 1]. A finite-difference check gives exactly these numbers. In symbols:Every linear layer's backward pass is this outer product, batched.
#Why Backprop Runs Backward
J_h = J_3 · J_2 · J_1. Matrix multiplication lets you group that product either way. Which end should you start from?[10, 1] · [[1, 1], [3, 2]] is one row times one matrix, and it gave [13, 12] directly.loss.backward() runs it.#Backprop, Derived End-to-End
Time for the payoff. We derive the gradient of the most common classification loss, softmax + cross-entropy on a one-layer network, using only the identities above. First, two definitions in plain words.
#Two Definitions You Need
Softmax turns a list of raw scores (called logits) into probabilities. It exponentiates each score so it is positive, then divides by the total so the results add up to 1. A bigger score gets a bigger share.
Cross-entropy loss for the true class is minus the log of the probability the model gave that class. If the model gave the true class probability 1, the loss is 0. The less probability it gave, the larger the loss. Entropy and KL divergence, the theory behind this loss, are covered properly in the Information Theory & Entropy lesson later in the track. Here you only need the one formula.
#Numbers First: A Three-Logit Trace
z = [2.0, 1.0, 0.1], and the true class is the second one, so the one-hot vector is y = [0, 1, 0].[7.3891, 2.7183, 1.1052], and these add up to 11.2125. Divide: p = [0.6590, 0.2424, 0.0986], which adds up to 1. The loss is −ln(0.2424) = 1.4170.p − y = [0.6590, −0.7576, 0.0986]. Test it by brute force. Add 0.001 to the first score and recompute the loss: it rises by 0.000659, which is 0.659 times 0.001. Add 0.001 to the second score: it changes by −0.000757. Add 0.001 to the third: it rises by 0.0000986. Central differences with a much smaller nudge give [0.6590, −0.7576, 0.0986], the same three numbers.Read the signs. The true class has a negative entry, meaning raising its score lowers the loss. The wrong classes have positive entries, meaning raising their scores raises the loss. The three entries add up to 0.
#Setup
x ∈ R^n, true class y ∈ {1, ..., k} (the one-hot vector y_one_hot ∈ R^k), parameters W ∈ R^{k × n} and b ∈ R^k.The forward pass:
∂L/∂W, ∂L/∂b, ∂L/∂x. Reverse-mode says: compute ∂L/∂z first, then propagate backward.#Step 1: ∂L/∂p
The loss only depends on one entry of p:
#Step 2: ∂p/∂z (Softmax Jacobian)
S = Σₖ e^{z_k}, so p_i = e^{z_i} / S. Use the quotient rule: (derivative of the top × bottom − top × derivative of the bottom) / bottom².e^{z_i} has derivative e^{z_i}, and S also has derivative e^{z_i}:∂p_i/∂z_i = (e^{z_i}·S − e^{z_i}·e^{z_i}) / S² = p_i − p_i² = p_i(1 − p_i).e^{z_i} does not depend on z_j, so its derivative is 0, while S has derivative e^{z_j}:∂p_i/∂z_j = (0·S − e^{z_i}·e^{z_j}) / S² = −p_i p_j.p₀(1 − p₀) = 0.6590 × 0.3410 = 0.2247, which matches the first diagonal entry of the numeric Jacobian. Both cases fit one formula:J_softmax = diag(p) − p pᵀ.#Step 3: ∂L/∂z: The Punchline
∂L/∂z_j = Σ_i (∂L/∂p_i)(∂p_i/∂z_j). Only the i = y term in the sum is nonzero:In vector form:
Bridge: The σ(x)(1−σ(x)) sigmoid derivative from the previous lesson is the chain rule's first link in every binary-classification backprop. The (p − y_one_hot) softmax+cross-entropy gradient derived here is the same idea generalised to multiclass.
#Step 4: Gradients to W, b, x
∂L/∂z in hand, the rest is the linear-layer identities from above:That is the entire derivation: a few lines of math, three identities, one cancellation. In a deep network the same pattern repeats at every linear layer, strung together by the chain rule.
To see the backward walk on a smaller graph, the viz below follows one neuron's loss back through its graph. Step the backward pass to watch each local derivative multiply into the gradient, then use the nudge button to compare the analytic gradient with the real change in the loss when you nudge a weight.
#Live: Compute a Jacobian Two Ways
jax.jacfwd computes analytically; the numerical path is the sanity check (torch.autograd.gradcheck) that ML libraries use to catch bugs in their backward rules.#Train a Softmax Classifier With Your Own Backward Pass
p − y, and the update rule are each a few lines. At the end it checks the hand gradient against finite differences.W and b at zero, every class gets probability 1/3, so the loss is exactly ln 3 = 1.0986. The printed loss at step 2 is 0.566, at step 5 it is 0.2601, at step 10 it is 0.1632, at step 20 it is 0.1086, and at step 50 it is 0.0667. Accuracy starts at 0.33, reaches 0.98 by step 2, and prints 1.0 from step 20 on. The hand-written gradient agrees with finite differences to about 1e-11 for dW and 5e-12 for db, and the loss after the final update is 0.066.p − y and the outer product, and a gradient check that catches mistakes. A framework does the same thing for every layer, automatically.#Common Misconceptions
Automatic Differentiation in Machine Learning: A Survey
Atılım Güneş Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, Jeffrey Mark Siskind (2018)
On the Variance of the Adaptive Learning Rate and Beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, Jiawei Han (2020)
#Try It Yourself
Every framework ships a gradient check, and it is the first tool to reach for when a hand-derived backward pass looks suspicious. The idea is simple: nudge each input a tiny amount in both directions, watch how the output moves, and compare that slope with your formula. Here you build the check yourself and use it on the softmax Jacobian and on the softmax + cross-entropy gradient from this lesson.
Tests · Verify numeric_jacobian returns a 4x4 matrix for softmax, np.allclose against diag(p) - p p^T is True with max abs error below 1e-8 (about 6e-11), and the cross-entropy gradient equals p - y = [0.7101, 0.0354, -0.8416, 0.0961] with max abs error below 1e-8.
True with a max absolute error of about 6e-11, and the stretch check prints the gradient p - y = [0.7101, 0.0354, -0.8416, 0.0961] for the true class at index 2, matching finite differences to about 1.6e-10. Those exact digits can shift slightly between machines, but the error should always sit many orders of magnitude below 1e-6.eps that is far too large, or a transposed Jacobian. Notice that the negative entry of p - y lands on the true class: the gradient pushes that logit up and the others down, which is the error-vector story from this lesson.#Putting It Together
Three concrete takeaways:
- Jacobian + chain rule = backprop. A neural network is a long composition of functions. Its gradient is one matrix product per layer, walked from output to input. Reverse-mode AD does this in time proportional to a single forward pass.
- Five identities, one convention.
∂(Ax)/∂x = A,∇(x^T y) = y,∇(x^T A x) = (A + A^T) x,∇‖x‖² = 2xand∇tr(AB) = B^Tlet you derive the gradient of many standard ML losses in minutes, and a gradient always has the shape of its variable. ∂L/∂z = p − y_one_hotis the gradient of softmax + cross-entropy, and it is the cleanest, most useful gradient in supervised learning. It says: the gradient is the error vector. You have now checked it by hand on three logits, derived it, and trained a classifier with it.
loss.backward(). The frameworks do the same steps for you at much larger scale.#Quick Check
For a function f: R^n → R^m, what is the shape of its Jacobian (one row per output)?