Checkpoint: Foundations Review
After this lesson, you will be able to:
- Find out which of the first fifteen lessons you really own and which only felt familiar while you read them
- Combine two or three ideas in one problem: a dot product and a perpendicularity test, a chain rule with numbers, a base rate and a payoff
- Compute a tiny linear model, its mean squared error and one gradient descent update by hand, then confirm it in code
- Summarise a small dataset with quartiles and an outlier fence, and say what a log transform does to it
- Turn a wrong answer into a repair plan: the lesson title, the section, and the reason it matters
Before You Start
#How to Use This Checkpoint
Take it closed-book. Close the earlier lessons, keep a sheet of paper and a calculator, and write down your answer to each question before you click anything. The quiz blocks reveal an explanation as soon as you answer, so clicking first teaches you nothing about what you know.
Fifteen questions sit in four short groups: toolkit and linear algebra, calculus and training, describing data, and probability. Then three longer problems each mix several lessons, and each ends with a code cell where you check your own working. Allow about 40 minutes in total, and do the groups in order.
Score yourself as you go, one point per question, and read the repair map at the end for what to do with your total.
#Group A: Toolkit and Linear Algebra
Four questions on logs, dot products, matrices and eigenvectors. Compute first, then choose.
Three independent events have probabilities 0.5, 0.25 and 0.125. A script stores the log base 2 of each probability and adds the three numbers. What does the script hold, and what is the product of the three probabilities?
#Group B: Calculus and Training
Four questions on the chain rule with numbers, a descent step, a diverging loss, and the shapes inside backprop.
A model predicts w x with w = 2 and x = 3, and the true value is y = 5. The loss is L = (w x - y)^2. What is dL/dw at this point?
#Group C: Describing Data
Three questions on reading a summary, what a log gap means, and what a density value is.
A box plot of response times in milliseconds reports: minimum 40, Q1 52, median 60, Q3 71, maximum 900. Which statement is correct?
#Group D: Probability
Four questions on counting, Bayes with a base rate, independence, and an expectation and variance.
A team has 6 engineers. The on-call rota names a first, second and third responder, in that order, with no repeats. A separate review panel is any 3 of the 6 with no roles. How many rotas and how many panels are there?
#Worked Problem 1: A Tiny Linear Model
This one mixes the Math Toolkit (summation), Optimization & Gradient Descent and Linear Regression: Your First Model. Try it on paper first.
A model predicts y-hat = w x + b. Three data points have x = 1, 2, 3 and true y = 3, 5, 6. Start from w = 1 and b = 0 with learning rate 0.1.
- Compute the three predictions and the three errors (prediction minus truth).
- Compute the mean squared error (MSE).
- Compute the two gradients, dL/dw = (2/3) times the sum of error x x, and dL/db = (2/3) times the sum of the errors.
- Take one update step and say whether the MSE fell.
Now check your working, and then try the stretch.
Tests · Verify predictions [1, 2, 3], errors [-2, -3, -3], MSE 7.3333, dw about -11.3333, db about -5.3333, then w about 2.1333, b about 0.5333 and a new MSE of about 0.3407. In the stretch, lr 0.1 gives a small MSE after 5 steps while lr 0.5 and 0.9 give enormous values.
The cell prints predictions [1. 2. 3.], errors [-2. -3. -3.] and an MSE of 7.3333. The gradients are dw = -11.3333 and db = -5.3333, one step gives w = 2.1333, b = 0.5333, and the MSE drops to 0.3407. In the stretch, a learning rate of 0.1 gets to an MSE of 0.2208 after five steps, but 0.5 reaches about 26,667,254 and 0.9 about 24 billion. The same data and the same formula diverge once the step is too big, which is exactly the diagnosis in Group B.
#Worked Problem 2: Reading a Small Dataset
This one mixes Descriptive Statistics, Quantiles and the Math Toolkit's logs. Nine page-session lengths, in minutes: 4, 5, 6, 7, 8, 9, 11, 14, 60.
- Compute the mean, the median and the sample standard deviation.
- Compute Q1, Q3, the IQR and the two outlier fences (the 1.5 x IQR rule), using the NumPy interpolation rule.
- Compute the z-score of 60 and say whether the usual rule of three flags it.
- Take log base 10 of every value. What happens to the mean, the median and the skew?
Tests · Verify mean 13.78, median 8.0, sd 17.61, Q1 6.0, Q3 11.0, IQR 5.0, fences -1.5 and 18.5 with only 60 flagged, largest z 2.63 not flagged by the rule of three, log10 mean 0.972 and median 0.903 with 9.37 minutes back on the original scale, a log upper fence of 1.436 that 60 (1.778) still exceeds, and skewness 2.33 raw against 1.44 on the log scale.
The cell prints a mean of 13.78, a median of 8.0 and a standard deviation of 17.61. Q1 is 6.0, Q3 is 11.0, the IQR is 5.0 and the fences are -1.5 and 18.5, so only 60 is flagged. Its z-score is 2.63, so the rule of three misses it. On the log scale the mean is 0.972 and the median 0.903, which is 9.37 minutes after converting back. The upper fence on the log scale is 1.436, and log10(60) is 1.778, so 60 is still an outlier. The skewness falls from 2.33 to 1.44. A log transform reduces the pull of a tail. It does not make a real extreme disappear, and deciding what to do with it is still your call.
#Worked Problem 3: A Screening Test and a Payoff
This one mixes Probability & Bayes' Theorem with Random Variables, Expectation & Variance. A condition affects 1% of people. A screening test has sensitivity 95% (positive for 95% of people who have the condition) and a false-positive rate of 8% (positive for 8% of people who do not). Every positive result triggers a follow-up that costs 30. Finding a true case is worth 400 (so the follow-up of a true case nets 400 - 30 = 370, and the follow-up of a healthy person nets -30).
- Someone tests positive. Compute P(positive) and then P(condition given positive).
- Let X be the net payoff of one positive result. Write its distribution, then compute E[X], Var(X) and the standard deviation.
- Out of 1,000 people screened, how many are flagged, and what is the expected total net payoff?
Tests · Verify P(positive) 0.0887 and posterior 0.1071, E[X] 12.84, Var(X) 15301.06 and sd 123.70, 88.7 flagged per 1000 with an expected net of about 1139, and a stretch posterior of about 0.49 when the false-positive rate drops to 0.01.
Run it and you will see P(positive) = 0.0887 and P(disease given positive) = 0.1071. The payoff has E[X] = 12.84, Var(X) = 15301.06 and a standard deviation of 123.70. Per 1,000 people, 88.7 are flagged and the expected net is about 1,139. The stretch prints 0.4897: cutting the false-positive rate from 8% to 1% lifts the posterior from about 11% to about 49%, which shows that the false-positive rate, not the sensitivity, was the lever here.
#The Repair Map
Add up your score out of 15 (one point per correct, uncorrected answer), then read the row that matches.
| Score | What it means | What to do |
|---|---|---|
| 13 to 15 | The foundations are solid | Go straight to Part 4. For any miss, read the one section named below |
| 10 to 12 | Mostly solid with a few patches | Reopen the sections named for each miss, redo those questions after a night's sleep, then continue |
| 6 to 9 | Real gaps, and they sit in specific places | Count the misses per group. Reread the lessons for the group with the most misses before moving on, because Part 4 leans on probability and statistics hardest |
| 0 to 5 | The material has not stuck yet | Do not push on. Redo the lessons in order, writing each worked example by hand, and retake this checkpoint before Part 4 |
For each wrong answer, here is what to revisit and the one-sentence reason it is worth ten minutes.
| Question | Topic | Lesson and section to revisit | Why it matters |
|---|---|---|---|
| 1 | Logs turn products into sums | Math Toolkit: Functions, Exponents, Logs & Sums, Why Machine Learning Lives on Logs | Every probabilistic loss is a sum of logs, and tiny products underflow |
| 2 | Dot product and perpendicularity | Vectors & Spaces, The Dot Product: Measuring Alignment | Similarity and attention scores are dot products, and zero is the only perpendicular value |
| 3 | Matrix times vector, determinant | Matrices as Transformations, Matrix-Vector Multiplication and The Determinant | A layer is a matrix, and the determinant says whether it squashes information away |
| 4 | Eigenvectors and eigenvalues | Eigenvalues & SVD, The Core Equation | PCA and stability analysis both ask which directions a matrix only stretches |
| 5 | Chain rule with numbers | Calculus: Derivatives & Gradients, The Chain Rule: Derivatives Through Composed Functions | Backprop is the chain rule applied layer by layer |
| 6 | One descent step | Optimization & Gradient Descent, The Algorithm: The Most Important Equation in Modern ML | Every training loop repeats this single subtraction |
| 7 | Diagnosing a diverging loss | Optimization & Gradient Descent, The Learning Rate: The Most Important Hyperparameter | A step that is too big is the most common reason a loss curve climbs |
| 8 | Gradient shapes in backprop | Vector & Matrix Calculus: Jacobian, Hessian & Backprop Rules, Backprop, Derived End-to-End | A gradient has the shape of its parameter, which catches most shape bugs |
| 9 | Quartiles, fences, robust statistics | Quantiles, Shape & Reading Real Data, Outliers: Detect First, Then Decide, and Robust Statistics | Choosing a summary that one bad row cannot move is basic data hygiene |
| 10 | What a log gap means | Math Toolkit, Logarithms: The Undo Button for Exponents | A step on a log scale is a multiplier, which is why logs tame long tails |
| 11 | Density versus probability | From Data to Distributions, Density Is Not Probability | Likelihoods and continuous models all use densities, and a height above 1 is fine |
| 12 | Ordered versus unordered counts | Counting & Probability Foundations, Counting Rules | Counting errors silently corrupt every probability built on them |
| 13 | Bayes with a base rate | Probability & Bayes' Theorem, The Medical Test: A Mind-Blowing Result | Spam filters, fraud alerts and diagnostics all depend on the base rate |
| 14 | Independence, union | Counting & Probability Foundations, Independent vs Mutually Exclusive | Models assume independence, and you should know how to test it |
| 15 | Expectation and variance of a table | Random Variables, Expectation & Variance, Expectation: The Long-Run Average and Variance: How Spread Out Are the Outcomes? | A loss function is an expectation, and variance is its uncertainty |
Worked Problem 1 also sends you to Linear Regression: Your First Model, the sections on the MSE loss and the update step, if your hand computation of the errors or gradients went wrong. Worked Problems 2 and 3 are covered by the rows for Questions 9, 10, 13 and 15.
#What You Can Now Do
If you scored well, or repaired what you missed, you can do all of these without looking anything up.
- Compute a dot product, test perpendicularity, and read a determinant or an eigenvalue from a small matrix
- Apply the chain rule with numbers, take a gradient descent step, and recognise a learning rate that is too large from the loss curve alone
- Work out the shape of a gradient in backprop and compute a tiny model's predictions, MSE and update by hand
- Summarise a dataset with robust tools, flag outliers with the IQR fence, and say what a log transform does and does not fix
- Use Bayes with a base rate, naming sensitivity and false-positive rate separately, and compute an expectation and variance from a small table
#Key Takeaways
- A checkpoint is a measuring instrument, so take it closed-book and mark any peeked answer wrong
- Mixed questions test whether you can choose the tool: a dot product of 4 is not perpendicular, because only exactly 0 is
- Training problems share one chain: the chain rule gives dL/dw = 6 for w = 2, x = 3, y = 5, and a step of 1.1 on w squared multiplies w by -1.2, so the loss climbs
- Small datasets teach robust habits: for 4, 5, 6, 7, 8, 9, 11, 14, 60 the IQR fence at 18.5 flags 60 while the z-score of 2.63 does not
- Bayes needs the base rate: sensitivity 90%, false-positive rate 5% and prevalence 2% give about 26.9%, not 90%
- The repair map turns each wrong answer into a named lesson and section, which is how a low score becomes a short, specific plan