Calculus: Derivatives & Gradients
After this lesson, you will be able to:
- Measure the slope of a curve by hand with a shrinking gap h, and watch the numbers settle on the derivative
- Spot the power rule (n x^(n-1)) in your own table of numbers, and use it with the constant and sum rules
- Compute partial derivatives (rate of change in one direction at a time) and combine them into a gradient that points uphill like a compass
- Use the chain rule (connecting rates of change step by step) on a two-step function, and check it against a finite-difference slope
Before You Start
#The Speedometer of Mathematics
You already understand derivatives — you just did not know it. Every time you checked your speedometer, watched a stock price ticker, or noticed your phone battery draining faster than usual, you were thinking about rates of change. This lesson just gives that intuition a name and a formula.
This seemingly simple idea — measuring how much the output wiggles when you wiggle the input — is the mathematical engine behind all of machine learning. Every time an AI model learns from data, it computes derivatives to figure out which direction to adjust its settings.
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
#Single-Variable Derivatives
A derivative is the slope of a curve at one point. You met slope in the Math Toolkit lesson as rise over run: how far the output goes up for each step you take across. On a straight line that number is the same everywhere. On a curve it changes from point to point, so we need a way to measure it at one spot.
#Measuring a Slope by Hand
Take the curve f(x) = x², which squares its input. Stand at x = 2 and take a step of width h = 1, to x = 3.
- At x = 2 the output is f(2) = 4.
- At x = 3 the output is f(3) = 9.
- The rise is 9 − 4 = 5 over a run of 1, so the slope is 5.
That slope is rough, because the curve bends during the long step. Shrink the step to h = 0.1. Now f(2.1) = 4.41, the rise is 0.41, and the slope is 0.41 / 0.1 = 4.1. A smaller step gives a better reading. The general recipe is (f(x + h) − f(x)) / h. Here it is for x = 1, 2 and 3 with four shrinking values of h, computed in Python:
| x | h = 1 | h = 0.1 | h = 0.01 | h = 0.001 |
|---|---|---|---|---|
| 1 | 3.0 | 2.1 | 2.01 | 2.001 |
| 2 | 5.0 | 4.1 | 4.01 | 4.001 |
| 3 | 7.0 | 6.1 | 6.01 | 6.001 |
Read each row. As h shrinks, the slope settles: row x = 1 on 2, row x = 2 on 4, row x = 3 on 6. Those are exactly 2 × x. So the slope of x² at x = 3 is 6, and the slope of x² at any x is 2x. That is a derivative, found from numbers alone.
The line through the two points you measured is called a secant. The picture below lets you shrink h and watch the secant line turn into the tangent line, the line that just touches the curve at one point.
Now give the idea its name. The number the slope settles toward as h shrinks is the derivative, written f'(x) and read "f prime of x". The short way to say "as h shrinks toward zero" is the symbol lim (for limit):
Below is a live tangent-line tracer. Drag the point along the curve and watch the slope (the derivative) update — that slope IS the derivative at that point.
#Spotting the Power Rule
The table showed that x² has slope 2x. Is that a one-off? Try x³ at x = 2, where f(2) = 8. The slopes for h = 1, 0.1, 0.01 and 0.001 are 19, 12.61, 12.0601 and 12.006, settling on 12. Now x⁴ at x = 1: the slopes are 15, 4.641, 4.0604 and 4.006, settling on 4.
Put the three results beside each other:
- x² has slope 2x, so at x = 3 it is 6.
- x³ has slope 3x², so at x = 2 it is 3 × 4 = 12.
- x⁴ has slope 4x³, so at x = 1 it is 4 × 1 = 4.
The pattern is the power rule: bring the exponent down in front, and lower the exponent by one. In symbols, the slope of xⁿ is n xⁿ⁻¹. For plain x (n = 1) it gives 1 × x⁰ = 1, which says a straight line y = x has slope 1 everywhere.
Two more rules complete what you need for the rest of this lesson:
- Constants vanish: the derivative of a constant is zero, because a constant never changes.
- Sum rule: the derivative of a sum is the sum of the derivatives. A multiplier on x just comes along, so 3x has slope 3.
Together they give the slope of f(x) = x² + 3x + 5 as 2x + 3. At x = 2 that is 7, and measuring with h = 0.001 gives 7.001. Same answer, no limit needed.
#Partial Derivatives: One Direction at a Time
Here is a function of two inputs, f(x, y) = x² + xy + y², and a concrete point (1, 2). Freeze y at 2. The function becomes f = x² + 2x + 4, which only depends on x, and its slope is 2x + 2. At x = 1 that is 4. Now freeze x at 1 instead. The function becomes f = 1 + y + y², with slope 1 + 2y. At y = 2 that is 5.
So at (1, 2) the slope going east (x direction) is 4 and the slope going north (y direction) is 5. The trick that did it: treat the other variable as a plain number. Here is the same idea as a formula:
The curly ∂ (read "partial") is just the derivative symbol telling you the function has several inputs.
#The Gradient: All Slopes at Once
The picture below shows the same function f = x² + xy + y². Read the two slopes at a point and watch them combine into one arrow, then spin the direction dial and see where the rate of climb peaks.
Why does the gradient point uphill as steeply as possible? Use the numbers at (1, 2). If you walk in a direction at angle θ from east, the slope you feel is 4 cos θ + 5 sin θ: the east slope times the east part of your step, plus the north slope times the north part. (Cosine and sine just turn the dial's angle into an east part and a north part.) Here is that slope for four dial settings:
| Direction | Angle | Slope you feel |
|---|---|---|
| due east | 0 degrees | 4.000 |
| between | 30 degrees | 5.964 |
| along the gradient [4, 5] | 51.34 degrees | 6.403 |
| due north | 90 degrees | 5.000 |
The rate peaks at 51.34 degrees, which is exactly the direction of [4, 5], and the peak value is the length of the gradient: √(4² + 5²) = √41 = 6.403. Every other direction feels a smaller slope, and the opposite direction feels −6.403, the steepest way down. So the gradient points in the direction of steepest ascent, and its length says how steep that ascent is.
At a local minimum of a function, what is the gradient?
When the gradient is the zero vector, the ground is flat in every direction. Optimizers use that as a signal that they have stopped moving.
#Visualize Gradients
Two pictures show gradients at work on a loss surface. In both, one rule applies: the gradient points uphill, and you step the opposite way.
First, the loss surface drawn as contour rings. Each ring joins points with the same loss, darker rings are lower, and the gradient at any point is perpendicular to the ring through it, pointing toward the higher rings. This picture draws a ball and its trail, not arrows. Pick a surface, press play, and watch the ball walk downhill, which is the direction of minus the gradient.
Watch the ball's steps. They get shorter as it nears the bottom, because the surface flattens there and the gradient shrinks.
Second, a top-down view called a gradient field, with an arrow at each grid point. These arrows are drawn pointing downhill, which is minus the gradient, because that is the direction you step. The gradient itself points the opposite way, uphill. Every arrow has the same length, so they show direction only, not steepness. Click the surface to drop a particle and watch it follow the arrows downhill.
For f(x, y) = 3x² + y², what is ∇f at the point (1, 2)?
#Step-by-Step: Computing a Gradient
#The Function
Let f(x, y) = x^2 + 2xy. We want to find the gradient at any point (x, y). This is a simple function, but the process works identically for functions of millions of variables.
#Partial Derivative with Respect to x
#Partial Derivative with Respect to y
#Assemble the Gradient
The gradient is the vector of both partial derivatives: nabla f = [2x + 2y, 2x]. At the point (1, 3), the gradient is [2(1) + 2(3), 2(1)] = [8, 2]. This vector points in the direction of steepest increase of f at (1, 3).
#Interpret the Result
The gradient [8, 2] at (1, 3) tells us: increasing x has 4 times the effect on f as increasing y at this point. If we wanted to decrease f, we would step in the direction [-8, -2] — opposite the gradient. That is exactly one step of gradient descent.
#The Chain Rule: Derivatives Through Composed Functions
Take y = (3x + 1)² at x = 2. Break it into two steps. The inner step is u = 3x + 1, and the outer step is y = u².
- Run the inner step: u = 3 × 2 + 1 = 7.
- Run the outer step: y = 7² = 49.
- Slope of the outer step at u = 7: the power rule gives 2u = 14.
- Slope of the inner step: u = 3x + 1 rises 3 for each 1 of x, so the slope is 3.
- Multiply the two slopes: 14 × 3 = 42.
Check by measuring. At x = 2.001 the inner step gives u = 7.003 and the outer step gives y = 49.042009. The rise is 0.042009 over a run of 0.001, a slope of 42.009, which is settling on 42. The chain rule works because a nudge of 1 in x moves u by 3, and each nudge of 1 in u moves y by 14, so a nudge of 1 in x moves y by 3 × 14 = 42. In symbols:
The general recipe has three steps: peel off the outermost layer, differentiate it, then multiply by the derivative of what is left inside. A neural network is a long chain of such steps, so the same recipe applies layer after layer. The math does not get harder with more layers; the product just gets longer.
Apply the chain rule: if y = (3x + 2)⁵, what is dy/dx?
#Preview: The Chain Rule in a Neural Network
A neural network is a chain of many layers, and the training method that applies the chain rule backward through all of them is called backpropagation. The Vector & Matrix Calculus lesson builds it properly. Here is a preview: run a forward pass, then watch the gradients flow back, each one multiplied by the local slope at its layer.
#Try It Yourself
You will measure slopes the way the table did, then check the power rule, the chain rule and the gradient numbers from this lesson. Each TODO is one step; the solution is below the starter code.
Tests · Verify the three rows settle on 2, 4 and 6 (last column 2.001, 4.001, 6.001); x^3 at 2 gives 12.0 measured and 12 from the power rule; the chain rule gives 42 and the measured slope is 42.0; the partials are 4.0 and 5.0 with gradient length 6.4031; and the slope at 30 degrees is 5.9641.
Running the solution prints these rows, which are the table from earlier in the lesson:
- x=1: [3.0, 2.1, 2.01, 2.001]
- x=2: [5.0, 4.1, 4.01, 4.001]
- x=3: [7.0, 6.1, 6.01, 6.001]
Then it prints 12.0 against 12 for the power rule, 42 against 42.0 for the chain rule, the partials 4.0 and 5.0, the steepest slope 6.4031, and 5.9641 for the slope at 30 degrees. Every number matches the hand work above.
Learning representations by back-propagating errors
David Rumelhart, Geoffrey Hinton, Ronald Williams (1986)
The paper that popularized backpropagation for training neural networks. Showed that the chain rule, applied systematically through a network, could learn useful internal representations. This single idea enabled the deep learning revolution.
#Symbolic Derivatives in Your Browser
Let's stop computing derivatives by hand and let SymPy do it. The playground below uses real symbolic differentiation — change the function and watch the derivative recompute. The first function is the x³ − 2x + 1 whose slopes you can check against the power rule.
#Key Takeaways
- Derivatives measure sensitivity. A derivative is the number the slope (f(x + h) − f(x)) / h settles toward as h shrinks. For x² it settles on 2x, the first case of the power rule n xⁿ⁻¹
- The gradient bundles all partial derivatives. For multi-variable functions, the gradient is a vector pointing in the direction of steepest ascent, and its length tells you how steep that ascent is (for f = x² + xy + y² at (1, 2) it is [4, 5], length 6.403)
- The chain rule multiplies slopes through composed steps. For y = (3x + 1)² at x = 2 the inner slope 3 times the outer slope 14 gives 42, and the same product grows longer in a neural network
- You step opposite the gradient to reduce a loss. The gradient points uphill, so gradient descent walks the other way, which the next lesson turns into an algorithm
- At a minimum, the gradient is zero. Optimizers read a gradient near zero as a sign they have stopped moving, though a peak or saddle is flat too
#Quick Check
What does the gradient of a function point toward?