Vectors, matrices, derivatives, probability, optimization, and stochastic calculus — from zero to the math you need to read modern ML papers. 31 lessons, foundations through advanced.
Functions, logs and sums first, then vectors, matrices and eigenvalues: the language models use to hold and transform data.
The few ideas every later lesson leans on: functions, exponents and e, logarithms (and why ML lives on them), sigma and pi notation, and how to read a formula word by word.
Understand vectors as geometric objects with direction and magnitude, and learn how vector spaces form the foundation of ML data representation.
See matrices as space-warping machines. Every neural network layer is a matrix multiplication — learn why that matters.
Discover the special directions that matrices stretch without rotating, and how SVD decomposes any transformation into simple pieces.
Derivatives and the gradient-descent loop, a first real model (linear and logistic regression), then tensors and backpropagation.
From single-variable derivatives to multivariable gradients — the mathematical engine behind backpropagation.
The algorithm that makes machines learn. Follow a ball rolling downhill through a loss landscape and understand learning rates, momentum, and convergence.
Fit a line, measure the miss with squared error, find the best line by calculus and by gradient descent, see overfitting, and turn a line into a probability with logistic regression.
Index notation, einsum semantics, batched ops, broadcasting, and computational graphs — the language every transformer and CNN forward pass speaks.
Jacobian, Hessian, key matrix-calculus identities, and a full end-to-end derivation of backprop for a softmax + cross-entropy classifier.
Summaries and shapes of real data, counting, probability and Bayes, random variables, then a checkpoint to test yourself before moving on.
Mean, median, mode, standard deviation, and the distributions that appear everywhere in ML.
Percentiles, quartiles, box plots, skewness, kurtosis, outliers and robust statistics, worked on one real-looking dataset.
How a histogram becomes a probability distribution: relative frequency, density, the empirical CDF, and why a sample is not the population.
Sample spaces, events, the probability axioms, permutations and combinations, independence, the law of total probability, and the chain rule, computed from a real table of counts.
Learn to think probabilistically. Bayes' theorem is the foundation of classification, generative models, and reasoning under uncertainty.
From random variables to expectation and variance — the statistical quantities that define every ML loss function and training process.
A self-test across the first fifteen lessons: mixed questions, three worked multi-step problems, and a repair map that tells you which lesson to revisit.
Correlation and causation, the named distributions, and how a sample speaks for a population.
Pearson and Spearman correlation, Anscombe's quartet, multicollinearity, confounders, and the paradox where a pooled table reverses every subgroup.
Uniform, Exponential, Gamma, Beta, Dirichlet, Student-t and Chi-square, plus heavy tails, with a decision guide for matching a distribution to your data.
Sampling distributions, the square-root law, sampling bias, and the bootstrap for any statistic.
Hypothesis tests, power and sample size, and Bayesian updating: how to decide whether a difference is real.
Understand hypothesis testing, p-values, confidence intervals, and the errors that lurk in every statistical decision.
Statistical power, effect size, A/B test sample sizes, peeking, Bonferroni, Holm and Benjamini-Hochberg, and winner's curse.
Prior, likelihood and posterior for a whole parameter: the Beta-Binomial update, credible intervals, Bayesian A/B tests, and Thompson sampling.
Entropy and cross-entropy, probability inequalities with Monte Carlo, the multivariate Gaussian, and maximum likelihood: where every loss function comes from.
Measure surprise and information content. Cross-entropy loss, KL divergence, mutual information, Jensen's inequality, and the ELBO that powers VAEs and diffusion.
Markov, Chebyshev, Hoeffding and Jensen inequalities, Monte Carlo estimation, importance sampling, and a first Metropolis sampler.
Joint, marginal, and conditional distributions; the multivariate Gaussian; conditioning via Schur complement; the reparameterization trick that powers VAEs and diffusion.
Where ML loss functions come from. Cross-entropy = MLE for categoricals; MSE = MLE for Gaussian noise; weight decay = Gaussian prior MAP. Fisher information and the natural gradient.
Markov chains and decision processes, convex optimization and stochastic calculus. Skip these if your goal is applied ML; they feed reinforcement learning and diffusion models.
Markov property, transition matrices, stationary distributions, and the Bellman equation — the mathematical foundation of RL, diffusion forward processes, and PageRank.
Convex sets and functions, Lagrangian duality, KKT conditions, and the SVM dual derivation — the optimization theory behind RLHF, TRPO/PPO trust regions, and constrained ML.
Brownian motion, Itô's lemma, SDEs, the score function, Langevin dynamics, denoising score matching, and the reverse SDE — the math powering diffusion models.
One dataset, the whole track, in two sessions: describe, test and quantify it; then model it, plan the follow-up and write the memo.
Describe the StreamBox dataset, look at its shape and relationships, quantify uncertainty with the bootstrap, test the plan gap, and start your analyst memo.
Fit a churn model by gradient descent, plan the follow-up experiment, take the Bayesian view, estimate by maximum likelihood, and write the final memo against a marking rubric.
10 interactive labs — hands-on exercises for this track
Fly over the terrain your optimizer must navigate — peaks are bad, valleys are good
Every neural network is just vectors being multiplied by matrices — build the intuition by dragging arrows on a coordinate plane.
Adjust μ, σ, n, p, and λ and watch the bell curve, bar chart, and shaded probability regions update live.
Warp a 2D plane with any 2×2 matrix and watch the eigenvectors stay fixed in direction — the axes the transformation preserves.
Drag a point along f(x) to watch its tangent line rotate and trace out f'(x) live. Zoom in until the curve BECOMES its tangent.
Trace df/dx backward through f(g(h(x))) — the mechanical pattern that is literally how backprop works.
Drag a slider and watch a straight line morph into a perfect replica of sin, cos, eˣ, ln(1+x), or 1/(1−x).
Click anywhere on a scalar field and watch a particle slide downhill along the negative gradient — always perpendicular to the contours.
Click any point on a 2D function. Gradient, Jacobian, Hessian computed live — eigenvalues classify the curvature as bowl/peak/saddle.
Click surfaces to drop particles. Convex = gradient descent always wins. Non-convex = 10 random inits show how initialization decides your fate.
425 questions across 17 modules — check how well you understood this track.