What’s one thing you learned? What’s still confusing?
Capstone Part 1: Describe, Test and Quantify StreamBox
Describe the StreamBox dataset, look at its shape and relationships, quantify uncertainty with the bootstrap, test the plan gap, and start your analyst memo.
Capstone Part 2: Model It, Plan It, Decide
Fit a churn model by gradient descent, plan the follow-up experiment, take the Bayesian view, estimate by maximum likelihood, and write the final memo against a marking rubric.
Welcome to AI
What is AI? What is machine learning? Start your journey here — no experience needed.
Interactive Labs for This Track
Loss Landscape
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Vectors & Matrix Operations
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Probability Distributions
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Ask questions, share insights
This lesson has two halves. Part 1 (core) tells the whole diffusion story with discrete steps, Gaussians and a gradient. Part 2 (Advanced, skippable) rewrites that story in continuous time with stochastic calculus. You can stop after Part 1 and still understand how diffusion models train and sample. Do Part 2 when you want to read Song et al. 2020 or any paper withdx_tin it.
Why split it this way? The only thing that genuinely needs Itô calculus is the step from "many small noise steps" to "one equation in continuous time". Everything a practitioner does with a diffusion model (train it, sample from it, debug it) can be explained without that step. So the lesson teaches the practitioner's story first and the physicist's language second.
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type the short code lines yourself.
x_0 = 2.0. A diffusion model destroys it slowly, one small step at a time, and each step does two things: shrink the value a little, then add a little Gaussian noise.β = 0.1 and one random noise draw ε = 0.5, one step gives:x_1 = √(1 − 0.1) · 2.0 + √0.1 · 0.5
= 0.9487 · 2.0 + 0.3162 · 0.5
= 1.8974 + 0.1581 = 2.0555
The value moved from 2.0 to 2.0555. Mostly the shrink factor 0.9487 acted, and the noise nudged it up a little. Repeat that many times and the original value fades while the noise piles up.
β_t at each step t:x_t depends only on x_{t-1}, with a Gaussian transition rule over continuous states.x_0 = 2. If you only add noise (no shrinking), the variance grows by β every step: after 10 steps of β = 0.1 it is 1.0, after 100 steps it is 10.0, and it never stops growing. With the shrink factor, the variance obeys v ← (1 − β) v + β, which settles at 1. If the data already has variance 1, the variance stays exactly 1 at every step, which is why this is called a variance-preserving chain. The destination is always N(0, 1), a distribution you know how to sample from.In DDPM, the forward step is x_t = √(1 − β_t) x_{t-1} + √(β_t) ε where ε ~ N(0, I). What does the contraction factor √(1 − β_t) accomplish?
α_t = 1 − β_t and ᾱ_t = α_1 · α_2 · … · α_t (the product so far, called alpha-bar). Then:β = 0.1 for 10 steps gives ᾱ = 0.9^10 = 0.3487, so the mean of x_10 is √0.3487 · 2.0 = 1.181 and its variance is 1 − 0.3487 = 0.651. The standard DDPM schedule (β rising from 0.0001 to 0.02 over 1000 steps) ends with ᾱ_1000 = 4.0e-5, so √ᾱ = 0.0064: the original data is scaled by 0.0064 and the rest is pure noise. At the halfway step ᾱ_500 = 0.0786.x_0 = 2 and compares it with the closed form.t, compute x_t from x_0 directly, and train on that one pair.p(x) ∝ exp(−x²/2). Take the log: log p(x) = −x²/2 + constant. Take the derivative: d/dx log p(x) = −x.x = 0.5 that is −0.5; at x = 3 it is −3; at x = 0 it is 0. The sign always points back toward the peak at 0, and the size is how far away you are. That derivative of the log-density is the score:x = 10 is −10. For the heavy-tailed Cauchy density p(x) = 1 / (π(1 + x²)) the score is −2x / (1 + x²), which at x = 3 is −0.6 and at x = 10 is only −0.198. Far from the data, how strongly a model pulls you back depends entirely on how fast the tails of the distribution fall off.Why does following the score undo noise? A noisy sample sits off to the side of where clean data lives. The score of the noisy data points from there toward where noisy points are denser, and that direction is also where the clean sample most likely came from. The next section makes this exact.
x_0 and corrupt it with Gaussian noise at some level, using the closed form from Part 1:x_t = α x_0 + σ ε, with α = √ᾱ_t, σ = √(1 − ᾱ_t), ε ~ N(0, I)x_t given a known x_0. Since x_t given x_0 is a Gaussian with mean α x_0 and variance σ², its score is∇_{x_t} log q(x_t | x_0) = −(x_t − α x_0) / σ² = −ε / σε yourself. The marginal score ∇ log p_t(x_t) is the one you actually want: the score of the noisy data over all possible x_0. You cannot compute that directly.s_θ(x_t) by least squares to predict the known target −ε/σ:s*(x_t) = E[−ε/σ | x_t] = E[∇ log q(x_t | x_0) | x_t]. The marginal density is p_t(x_t) = ∫ q(x_t | x_0) p(x_0) dx_0. Differentiate it under the integral and divide by p_t:∇ log p_t(x_t) = E[∇ log q(x_t | x_0) | x_t]−ε/σ therefore yields exactly the score of the noisy marginal. That is Vincent's result (2011), and it is why diffusion training looks like ordinary supervised learning.∇ log q = −(x_t − α x_0) / σ² inside the expectation:∇ log p_t(x_t) = −(x_t − α E[x_0 | x_t]) / σ²Solve for the clean-sample estimate:
ε instead of the score. They are the same thing in two outfits: s_θ(x_t) = −ε_θ(x_t) / σ. Training on ‖ε − ε_θ(x_t, t)‖² is denoising score matching with a rescaled target, and the single network is told t so it can handle every noise level.−2, half near +2, each bump with standard deviation 0.5. After corrupting at ᾱ = 0.5 (so α = σ = 0.7071), the noisy data is again two Gaussian bumps, so its true score has a closed form. That lets us compare a learned score against the truth.The cell below does denoising score matching with a tiny model (13 bump-shaped features, fitted by least squares) and then checks Tweedie's formula against a brute-force average.
x_t = −0.5 the learned value is −1.092 against the exact −1.036, and at x_t = −1.0 it is −0.578 against −0.614. The mean absolute gap over [−3, 3] is 0.0724. The exact score at x_t = 0 is 0, because 0 is the dip between the two bumps, where the two pulls cancel (the learned value is −0.045). Notice that the sign always points toward the nearer bump. The Tweedie line also checks out: at x_t = −1 the brute-force average of the clean samples is −1.845 against the formula's −1.849, and at x_t = 1 it is 1.856 against 1.849. The noisy point −1 most likely came from the bump at −2.Now the generation half. Starting from pure noise, run the chain backwards. One reverse step takes the score and does the Tweedie-style move plus fresh noise:
The cell below uses the exact score of the noised data at every step (a trained network would supply an estimate of it instead) and samples 100,000 points from pure noise.
ᾱ of 1.9e-05, so the chain starts from essentially pure noise. About 50.1% of the samples land right of zero. The right bump has mean 1.992 and standard deviation 0.501, and the left bump has mean −1.998 and standard deviation 0.501. The data had centres at ±2 and standard deviation 0.5, so the reverse chain recovered both the locations and the widths from pure noise.That is the entire diffusion model: a forward noising chain, a score (or noise) predictor trained by regression, and a reverse chain that follows the score with fresh noise. Part 1 ends here.
dW_t like a Gaussian random variable with variance dt" gets you most of the way, and this half teaches that heuristic, shows it in numbers, and ends by writing DDPM as one of these equations.You already know random walks from the Markov chain lesson. Here is the punch line of stochastic calculus: take a random walk, shrink the step size, speed up time, and you get Brownian motion.
dW_t engine produce everything from pure noise to drifting random walks to Geometric Brownian Motion (GBM), a multiplicative version used in finance, for instance behind the Black-Scholes option-pricing formula, which you do not need here.W_t satisfies:W_0 = 0 (start at the origin)s < t, W_t − W_s ~ N(0, t − s), and increments over disjoint time intervals are independentW_t is nowhere differentiableW_t, you see more wiggles, not a tangent line. The path has fractal-like structure all the way down. This is why ordinary calculus fails: there is no dW_t/dt to take a derivative of.A standard Brownian motion W_t has Var(W_t) = t. You simulate W_T over the interval [0, 4] and over [0, 16]. How does the typical magnitude of W_T compare between the two simulations?
f(t) against the noise dW_t. Discretise time into steps t_0 < t_1 < ... < t_n = T and define:f(t_i) independent of W_{t_{i+1}} − W_{t_i}. That independence yields clean martingale theory. (A martingale is a process with no predictable trend: your best forecast of its future value is its value now, like the running total in a fair coin-toss game.) Using the midpoint gives you the Stratonovich integral, which obeys ordinary chain-rule calculus but loses some of Itô's nice properties. ML almost universally uses Itô.dW_t live in different function spaces), but the discrete heuristic above is the only thing you need for ML papers.X_t evolves according to dX_t = μ dt + σ dW_t, and you want to track Y_t = g(t, X_t), some smooth function of time and the noisy state. In ordinary calculus the chain rule gives dY = (∂g/∂t) dt + (∂g/∂x) dX. In stochastic calculus you get an extra term:(dW_t)² ≈ dt (their variance is dt), not zero. In a Taylor expansion of g(X + dX) around X, the (dX)² term, which would normally be discarded as "higher order", contributes a non-negligible σ² dt. That is the Itô correction.(dW)² ≈ dt is a claim. Test it on the simplest case, Y = W². The ordinary chain rule says d(W²) = 2W dW. Itô says d(W²) = 2W dW + dt. Add up both sides from time 0 to T (and W_0 = 0):W_T² = 2 ∫ W dW + T∫ W dW has average zero: each term in the sum is W at the left end times a later increment, two independent pieces, and the increment has mean zero. So E[2 ∫ W dW] = 0. We also know E[W_T²] = Var(W_T) = T. The ordinary chain rule would predict E[W_T²] = E[2 ∫ W dW] = 0, which contradicts Var(W_T) = T. The extra dt is exactly what fixes the books.The cell simulates 20,000 Brownian paths with 1,000 steps each and prints every piece.
E[W_T²] = 0.9960 against the theory 1, an average of the Itô integral term E[2 ∫ W dW] = −0.0035 (standard error 0.0101, so consistent with 0), and an average of the squared increments E[Σ(dW)²] = 0.9995. So −0.0035 + 0.9995 = 0.9960: the whole of E[W_T²] comes from the squared increments, not from the "chain rule" term. The pathwise identity W_T² = 2 ∫ W dW + Σ(dW)² holds to 3.8e-14 on every path. And the sum of squared increments hardly varies from path to path (standard deviation 0.0449 around 1), which is why replacing (dW)² by its average dt is a sound move: it is almost not random at all.(Stratonovich calculus does obey the ordinary chain rule, at the cost of the integrand and the noise becoming statistically dependent. Some physics literature uses it; ML almost never does.)
What is the practical difference between the Itô and Stratonovich interpretations of a stochastic integral for ML purposes?
dx = -θ x dt + σ dW. It is the "spring plus random kick" model used in physics, neuroscience and finance. Watch the path get yanked back to zero whenever the noise sends it too far.Now run the same Euler-Maruyama update yourself: switch between the four models, drag the time cursor, and watch the histogram of X_t beside the paths settle onto the exact density. The diffusion forward process is the one from Part 1: it melts a two-bump "data" distribution into N(0, 1).
U(x) (a landscape in which a small U means a good place to be):β = 1 and U = −log p. Then the drift is ∇log p(x), the score, and the noise is √2 dW. Keep that picture in mind: following the score plus a steady jitter is the Langevin SDE.p_t(x) of the state at time t. For an SDE dx = f(x,t) dt + g(t) dW, the density evolves according to the Fokker-Planck equation:dx = ∇log p(x) dt + √2 dW, the stationary distribution (the t → ∞ limit of p_t) is exactly p(x). This is the foundation of Langevin sampling. Here is the check, in one dimension, step by step.∂p/∂t = −∂(f p)/∂x + (g²/2) ∂²p/∂x². Group the two terms as one: ∂p/∂t = −∂J/∂x, whereJ = f p − (g²/2) ∂p/∂xx. If J = 0 everywhere, nothing flows, ∂p/∂t = 0, and the density is stationary.f = ∂log p/∂x = p′/p and g² = 2. Then J = (p′/p) · p − (2/2) p′ = p′ − p′ = 0. The density p is stationary, whatever p is.f = −θx, g = σ. Try p ∝ exp(−θx²/σ²), a Gaussian with variance σ²/(2θ). Its derivative is p′ = −(2θx/σ²) p. Then J = −θx p − (σ²/2)(−2θx/σ²) p = −θx p + θx p = 0. ✓ So the stationary variance σ²/(2θ) seen in the simulation above follows from the equation, not just from simulation. The drift pulls mass inward at rate θ x p, and diffusion pushes it outward at exactly the same rate.J on a grid for the claimed stationary variance and for a wrong one, so you can see that a zero current is a real test.max |J| = 4.59e-06 for the claimed variance 0.32 (zero up to finite-difference error, against a drift flow of size 0.242), and max |J| = 8.07e-02 for the wrong variance 0.48. Only the right variance balances drift against diffusion.Consider the SDE dx = μ dt + σ dW (constant drift, constant diffusion). What kind of process is this?
dx = ∇log p(x) dt + √2 dW has stationary distribution p(x). Discretising this SDE with step size ε gives Langevin dynamics, the universal sampler:dx = f(x, t) dt + g(t) dW, the reverse-time SDE is:x_t = √(1 − β_t) x_{t-1} + √β_t ε. For a small β_t, √(1 − β_t) ≈ 1 − β_t/2, sox_t ≈ x_{t-1} − ½ β_t x_{t-1} + √β_t εThat is one Euler-Maruyama step of the SDE
(x_t + β_t s)/√α_t + √β_t z ≈ x_t + ½ β_t x_t + β_t s + √β_t z, is one Euler-Maruyama step of the reverse SDE with this f and g: the drift f − g² s = −½ β x − β s is applied backward in time. So DDPM is a discretisation of the variance-preserving SDE, and DDPM sampling is a discretisation of its reverse. Reopen the SDE simulator above, pick the diffusion forward process, and you are watching the forward half of this equation.β(t) rising from 0.1 to 20 over t in [0, 1], 1000 reverse steps and 100,000 samples.a(1) = 0.0066, 50.1% of samples on the right, a right bump with mean 2.003 and standard deviation 0.503, and a left bump with mean −2.001 and standard deviation 0.502. The continuous-time sampler recovers the same two bumps as the discrete one in Part 1, from pure noise.This is the unifying view: DDPM, score-based models, EDM and latent diffusion are all discretisations of the reverse SDE, with some way of parameterising the score. The architecture (U-Net, DiT) is implementation detail; the math is the same.
v_θ(x, t) such that integrating dx/dt = v_θ(x, t) from noise to data produces samples. It is the same idea (transport mass from a known distribution to data) with less stochasticity and often faster sampling. See the Flow Matching lesson in the Generative AI track for implementation details. The math primitives are the same as in this lesson, minus the dW term.Score-Based Generative Modeling Through Stochastic Differential Equations
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, Ben Poole (2021)
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel (2020)
Reading about Euler-Maruyama is not the same as writing it. Here you build the solver yourself and point it at the Ornstein-Uhlenbeck process from earlier. Then you measure the discretisation error two ways, so that it shows up clearly instead of drowning in sampling noise. First you compute the error of the Euler scheme exactly, with no randomness at all, by recursion. Then you run a large simulation and check that it lands on the Euler number rather than on the exact one.
Tests · Verify euler_maruyama returns shape (n_paths, n_steps + 1); the Euler-recursion mean error at t = 1 and variance error at T = 6 both shrink steadily as dt falls through 0.2, 0.1, 0.01, 0.001; the dt = 0.1 sample variance lies within a few standard errors of the Euler recursion value 0.2306 but many standard errors from the exact 0.2133; the GBM mean is near 1.105 and its median near 1.057.
t = 1 is 0.4463 and the exact stationary variance is 0.2133. The Euler recursion has no sampling noise, so its errors show the discretisation error alone. The mean error falls from 0.1101 at dt = 0.2 to 0.0525 at 0.1, 0.0050 at 0.01 and 0.0005 at 0.001. The variance error falls from 0.0376 to 0.0173, 0.0016 and 0.0002. Cutting dt by ten cuts the error by about ten: the error is first order in dt. At dt = 0.1 the Euler scheme settles on a variance of 0.2306, about 8% too large, because each step slightly over-counts the noise.dt = 0.1 (seed 42) the sample variance is 0.2323 with a standard error of 0.0015. That is only 1.1 standard errors from the Euler recursion value 0.2306 but 12.9 standard errors from the exact 0.2133. So the simulation reproduces the scheme faithfully, and the scheme itself is biased by the step size. This is the fair way to separate bias from noise: noise shrinks with more paths, bias shrinks with a smaller dt.For geometric Brownian motion over T = 1, one run printed a mean of 1.1018 against the exact 1.1052, and a median of 1.0550 against the exact 1.0565. The mean tracks exp(0.1) while the median sits lower by exp(-0.045), because the volatility drag pulls the typical path below the average one.
This lesson tied together earlier chapters of the math track:
q(x_t | x_{t-1}).q(x_t | x_0) and the reparameterisation trick (write a sample as mean plus scale times fixed noise) that made everything trainable.If you can derive the DDPM forward closed form from the Markov definition, state why regressing onto −ε/σ learns the score, and (for Part 2) recognise the Itô correction in a derivation and write down the reverse SDE, you have the math required to read any diffusion or score-based paper from 2020 onwards.
In the forward process x_t = √(1 − β) x_{t-1} + √β ε with β = 0.1, you start from x_0 = 2 and run 10 steps (so ᾱ = 0.9^10 = 0.3487). What is the distribution of x_10?