Linear Regression: Your First Model
After this lesson, you will be able to:
- Read a line y = w x + b as a model: say what the slope and intercept mean, make a prediction, and compute a residual (the miss on one point)
- Compute the mean squared error (MSE) of a line by hand, and explain why squares are used instead of absolute values
- Derive the two partial derivatives of the MSE and set them to zero to get the closed-form least-squares line
- Run gradient descent on the same loss from scratch in numpy, and explain why the learning rate depends on the scale of x
- Judge a fit with R-squared and residuals, and test it against an outlier
- Explain overfitting with a train and test error, and extend the loop to logistic regression with the cross-entropy gradient (p minus y) times x
Before You Start
#A Line as a Model
Here are five toy points: x is hours of practice in a week and y is problems solved. The data is made up, but every number below is computed. Try the guess line with w = 1 and b = 1.
| x | y (true) | y-hat = 1x + 1 | residual = y minus y-hat |
|---|---|---|---|
| 1 | 3 | 2 | 1 |
| 2 | 5 | 3 | 2 |
| 3 | 4 | 4 | 0 |
| 4 | 7 | 5 | 2 |
| 5 | 9 | 6 | 3 |
#Why Squares: The Loss
Turn each miss into a cost, then average the costs. The usual cost of a miss is its square.
Now try a better guess, w = 1.5 and b = 1. Its predictions are 2.5, 4, 5.5, 7 and 8.5, so its misses are 0.5, 1, minus 1.5, 0 and 0.5. The squares are 0.25, 1, 2.25, 0 and 0.25, which add to 3.75, so the MSE is 0.75. The loss fell from 3.6 to 0.75, and that is what "a better line" means: a smaller loss.
Drag the sliders on the first tab of the widget below and watch the squares on each miss grow and shrink; the best line is the one with the least total square area.
Why squares and not absolute values?
You could average the absolute misses instead. For the guess line that gives 1.6, and for the better line 0.7. Squares are preferred for three reasons.
- A square is smooth, with no sharp corner at zero, so its derivative exists everywhere. Calculus and gradient descent both need that.
- Squaring punishes big misses much more. One miss of 3 costs 9, while three misses of 1 cost only 3 in total, so the line cannot ignore its worst point.
- The slope of the square is 2 times the miss, so near the best line the push gets gentler on its own. The absolute value pushes just as hard however close you are.
#The Best Line by Calculus
For a given set of points, L is a function of two numbers, w and b. Every pair gives a different line and a different loss. The lowest point of that bowl is the best line, and from the Derivatives lesson you know how to find a lowest point: slope zero in every direction. So take the two partial derivatives and set both to zero.
#Step 1: Name the miss
Let the prediction miss on point i be the prediction minus the truth, written m_i = (w x_i + b) minus y_i. The loss is then the average of m_i squared. Writing it this way flips the sign of the residual, which makes no difference once it is squared.
#Step 2: Differentiate one square
By the chain rule, the derivative of m squared is 2 m times the derivative of m. Here m = w x + b minus y, so the derivative of m with respect to w is x, and with respect to b is 1.
#Step 3: Average over the points
The derivative of an average is the average of the derivatives. That gives the two partial derivatives, one for each knob, and together they are the gradient of the loss.
#Step 4: Set both to zero and solve
The b equation says the average miss is zero, which gives b = mean of y minus w times mean of x. Putting that into the w equation leaves one equation in w alone, and its solution is the slope below.
On the five points the best line is y = 1.4 x + 1.4. Its misses are 0.2, 0.8, minus 1.6, 0 and 0.6, its squared misses add to 3.6, and its MSE is 0.72. That beats both guesses (3.6 and 0.75), and the gradient there is zero. Run the cell to check this and to fit the 12 points of hours studied against exam score that the widget uses.
For the 12 hours-versus-score points, the covariance of hours and score is 25.4306 and the variance of hours is 5.4306. Their ratio is the slope, 4.682864, and the intercept is 45.290921. Read that as: each extra hour of study goes with about 4.7 more points, and a student who studied zero hours is predicted to score about 45. Treat the second sentence with care: no student in the data studied zero hours, so the intercept is a rule for drawing the line, not a measured fact. The MSE of this line is 5.9885.
Two lines are fitted to the same five points. Line A has MSE 3.6 and line B has MSE 0.75. What can you conclude?
#Gradient Descent Finds the Same Line
The closed form is perfect for one line, but it does not carry over to a neural network, where no such formula exists. Gradient descent does carry over, so check that it finds the same answer. You already have the update rule from the last lesson, where eta is the learning rate. Now fill in the two derivatives you just computed.
The second tab of the same widget (back in the Why Squares section) runs exactly this loop on the 12 points: it starts at w = 0 and b = 0, draws the line moving, the loss curve falling and the path across a map of the loss. The cell below does it in plain numpy, and also shows what happens with the learning rate.
Three things stand out in the output.
- On the raw hours scale a safe rate is 0.03, and even then the loop needs 1,102 steps to get the gradient below 0.00001. The slope overshoots badly at first: after one step it is 18.58 against a true value of 4.68.
- A learning rate of 0.045, only 50 percent bigger, diverges at step 28. The loss grows from 4,426 to over a billion. There is a hard ceiling near 0.04 on this data.
- If you first standardise x, meaning subtract its mean (4.33) and divide by its standard deviation (2.33), a rate of 0.3 converges in 18 steps, and translating the answer back gives slope 4.682864 and intercept 45.290918, the closed-form line to five decimal places.
Why such a difference? The loss bowl is stretched. Raw hours run from 1 to 8.5 and their average is far from zero, so a small change in b and a small change in w have very different effects, and the bowl is a long narrow valley. A single learning rate must be small enough not to bounce off the steep wall, which makes the walk down the gentle floor crawl. Standardising x makes the bowl round, so the same rate suits both directions. The ratio of the steepest to the gentlest curvature on the raw hours scale is about 115, and on the standardised scale it is 1.
#How Good Is the Line?
Finding a line is easy. Knowing whether to trust it is the real skill. Two tools help.
Now a prediction. Add one more student to the 12: they studied 9.5 hours and scored 30.
One new point (9.5 hours, score 30) is added to the 12 points, and the least-squares line is refitted. The slope was 4.68. What happens to it?
For the 12 points the SSE is 71.86 and the SST is 1,500.92, so R-squared is 1 minus 0.0479, which is 0.9521. The residuals add to zero, as promised. After adding the outlier the slope drops from 4.68 to 1.51 and R-squared collapses from 0.95 to 0.08. One point out of 13 rewrote the whole story, and a single number like R-squared would have flagged it only if you looked.
The last lines of the cell fit a line to a perfect curve, y = x squared. R-squared is a respectable 0.9284, but the residual signs run plus, plus, plus, then nine minus and plus again, a U-shape that a plotted residual graph would show instantly. High R-squared with patterned residuals means the model is the wrong shape.
#Overfitting and the Train-Test Split
In the cell below the truth really is a straight line, y = 2 + 3x, but each observation has noise with standard deviation 0.5. We fit polynomials of degree 1, 3 and 9 to 10 noisy training points, then score them on 200 fresh points.
| Degree | Train MSE | Test MSE |
|---|---|---|
| 1 | 0.6617 | 0.2674 |
| 3 | 0.3954 | 0.4947 |
| 9 | 0.0000 | 1.5268 |
The training error falls steadily to zero, since degree 9 passes through every training point. The test error rises. The unavoidable noise alone gives a test MSE of 0.25, and the line gets close to that (0.2674) while degree 9 is six times worse. Averaged over 200 different training sets the pattern holds: 0.314, 0.354 and 1.985.
- Bias is the error from a model too simple to follow the real pattern. A line fitted to the curve y = x squared has high bias, which is what the patterned residuals showed.
- Variance is the error from a model so flexible that its shape changes a lot with each new batch of noisy training data. Degree 9 has high variance.
A good model balances the two, and the test error is how you find the balance. Never use the training error to choose between models: it can only go down as you add flexibility.
Model A gets train MSE 0.00 and test MSE 1.53. Model B gets train MSE 0.66 and test MSE 0.27. Which model do you expect to do better on new data, and why?
#From a Line to a Probability: Logistic Regression
So far y was a number. What if the answer is yes or no, such as whether a StreamBox user will churn (leave the app)? You could code stayed as 0 and churned as 1 and fit a straight line to sessions per week. The cell below tries it, and the trouble is that a line is unbounded: on the 200 users the fitted line (slope minus 0.1000, intercept 0.5858) gives 47 users a "probability" below 0, the lowest being minus 0.614. A probability must live between 0 and 1.
Now the lovely part. Push the derivative through the sigmoid with the chain rule, and the gradient of the average cross-entropy comes out in the same shape as before.
Do not take that on trust: the cell below checks the formula against finite differences, nudging w and b by a tiny amount and watching the loss. It then runs the same gradient-descent loop on the 200 StreamBox users. The cell starts by pasting the course's StreamBox generator.
The formula and the finite differences agree to six decimals (0.756021 and 0.150035). Training starts at loss 0.6931 and settles at 0.3005 near w = minus 1.2303 and b = 1.7300. Every extra weekly session lowers the score by 1.23, so a user with 0 sessions has a 0.849 chance of churning, 1 session 0.622, 2 sessions 0.325, 3 sessions 0.123 and 5 sessions only 0.012. Predicting churn whenever p is at least 0.5 gives an accuracy of 0.865.
Is 86.5% good? Compare it with the laziest possible model. 44 of the 200 users churned, so a model that says everyone stayed is right 78 percent of the time without learning anything. Ours beats that by 8.5 points: it catches 30 of the 44 churners, with 13 false alarms. Tab 3 of the same widget draws a related line through the same users, session minutes against sessions per week, so you can see how strongly the two behaviour columns move together.
#This Is the Whole Training Loop
Look back at what you did twice, once for a line and once for a probability.
- Model: a function with adjustable numbers, y-hat = w x + b, or p = sigmoid(w x + b).
- Loss: one number for how wrong the model is, the MSE or the cross-entropy.
- Gradient: the slope of the loss in every parameter direction, from the chain rule.
- Update: nudge each parameter downhill, w becomes w minus eta times its slope, then repeat.
A neural network with a billion weights runs exactly this loop. It has a bigger model made of many layers, usually a different loss, and an automatic way of computing the gradient (backpropagation, which the Tensor Notation, Einsum & Computational Graphs lesson and the Vector & Matrix Calculus lesson build up to). The four steps do not change, and neither do the practical lessons: put the inputs on a sensible scale, tune the learning rate, and always judge the model on data it has not seen.
#Try It Yourself
Here is a fresh dataset of eight made-up points: weeks since launch (1 to 8) against daily signups (14, 19, 21, 28, 30, 36, 37, 44). You will fit it twice and confirm that the two methods agree.
Tests · Verify the closed form gives slope 4.130952, intercept 10.035714 and MSE 1.394345; gradient descent on standardised x with learning rate 0.1 stops after about 100 steps and agrees with the closed form to within 1e-6; R-squared is about 0.9847 and the week 9 prediction is about 47.214. For the stretch, the two-feature logistic model reaches an accuracy of about 0.89 against 0.865 for sessions per week alone.
The solution prints a closed-form slope of 4.130952, intercept of 10.035714 and an MSE of 1.394345. Gradient descent on the standardised weeks takes 102 steps and lands on the same slope and intercept, with a largest difference of about 1.6e-09. R-squared is 0.9847, and the week 9 prediction is 47.214 signups (an extrapolation, so treat it as a guess). In the stretch, adding session length as a second feature lifts the churn accuracy from 0.865 to 0.89, with a final loss of 0.2133.
#Key Takeaways
- A line y-hat = w x + b is a model, the slope and intercept are its parameters, and a residual is the true value minus the prediction
- The MSE averages squared misses, so on the five toy points the guess line scores 3.6, a better one 0.75 and the best line, y = 1.4x + 1.4, scores 0.72
- Setting the two partial derivatives of the MSE to zero gives slope = cov(x, y) / var(x) and an intercept that passes through the averages; on the hours data that is 4.682864 and 45.290921
- Gradient descent reaches the same line: 1,102 steps at a safe rate on raw hours, divergence at 0.045, and 18 steps once x is standardised
- R-squared is 1 minus SSE over SST and the residuals show the shape; one outlier moved the slope from 4.68 to 1.51 and R-squared from 0.95 to 0.08
- Training error only falls as a model gets more flexible, so choose with test error: degree 9 had train MSE 0.0000 and test MSE 1.5268 against 0.2674 for a line
- Logistic regression passes the line through a sigmoid, uses cross-entropy, and has the gradient (p minus y) times x; it scored 86.5 percent on StreamBox churn against a 78 percent baseline
- Model, loss, gradient, update is the whole training loop that every neural network repeats
#Quick Check
A line y-hat = 2x is tried on the points (1, 3), (2, 5) and (3, 4). What is its mean squared error?