Capstone Part 2: Model It, Plan It, Decide
After this lesson, you will be able to:
- Write a logistic regression for churn as a feature matrix X, a weight vector w and a matrix-vector product
- Derive and verify the gradient of the cross-entropy loss, then train the model with a hand-written gradient-descent loop
- Compare a model against the 'always stayed' baseline and say why a strong fit here is partly an artefact
- Size a follow-up experiment, including the cost of watching several metrics
- Answer the same comparison with Beta posteriors and score predictions with entropy and cross-entropy
- Write and mark a final analyst memo against a four-criterion rubric
Before You Start
#Where Part 1 Left Off
You have the same 200 StreamBox users: plan, device, churned, session_min and sessions_per_week, plus a population of 10,000 session lengths. In Part 1 you found 44 of 200 churned (22.0%), Free churned 34.0% against 10.0% for Pro, and the gap was real but not proof of cause. The product manager's last question is still open: what should we do about it?
users from the first.#Milestone 6: Model Churn by Hand
You know the plan gap. Can a model predict who churns? The tool is logistic regression, which you met with one feature in Linear Regression: Your First Model. Here we use three features and write it with matrices.
One tiny example first
Take three users and one feature plus a constant 1 for the intercept (the bias). Stack them as rows of a matrix X and the weights as a vector w:
X = [[1, 0], [1, 1], [1, -1]], w = [-1, 2], true labels y = [0, 1, 0].
The matrix-vector product X w gives one score per user: [-1, 1, -3]. The sigmoid turns each score into a probability: 0.2689, 0.7311 and 0.0474. The cross-entropy loss for each user is minus the log of the probability put on what really happened, 0.3133, 0.3133 and 0.0486, and the average is 0.2250.
The gradient says which way to move w. The prediction errors p minus y are [0.2689, -0.2689, 0.0474], and X transposed times that vector, divided by 3, is [0.0158, -0.1055]. One gradient-descent step with learning rate 0.5 gives w = [-1.0079, 2.0527], and the loss falls from 0.2250 to 0.2194.
Now the names
Why is the gradient that simple? The derivative of the loss with respect to one user's score works out to p minus y, which is the Matrix Calculus and Backprop result for a sigmoid followed by cross-entropy. The chain rule then sends that error back to each weight through the matching column of X, and doing it for all columns at once is the product with X transposed. The cell does not ask you to trust this: it checks the formula against finite differences, the slope estimated by nudging each weight by a tiny amount.
For StreamBox the feature columns are sessions_per_week and session_min, each standardised (subtract the mean, divide by the standard deviation, so a typical value is 0 and one standard deviation is 1), plus plan as 0 for Free and 1 for Pro, plus a column of ones. Standardising puts the features on the same scale so a single learning rate suits all of them.
Before any training, all weights are zero, so every user gets the same score of 0 and the same probability. What is the average cross-entropy loss at that starting point?
The feature matrix has shape 200 by 4. The finite-difference check agrees with the analytic gradient to 5.5e-11, so the formula is right. Training starts at a loss of 0.6931 and falls to 0.6046 after 1 step, 0.3556 after 10, 0.2390 after 50, 0.2150 after 100 and 0.1950 after 500, where the curve has flattened. The learned weights are [-3.062, -3.200, -2.116, -1.607] for bias, sessions per week, session length and Pro.
All three feature weights are negative: more sessions, longer sessions and being on Pro each lower the predicted chance of churn. Accuracy, using 0.5 as the cutoff, is 92.0%. The always-stayed baseline, which predicts that nobody churns, is right for 156 of 200 users, 78.0%. The model's average cross-entropy is 0.1950 nats, or 0.2813 bits.
Beating a 78% baseline by 14 points looks impressive, and it is partly an artefact. The session columns were generated from churn, so the model is mostly reading the answer off them. A model built only on plan scores no better than the baseline, because even the Free plan churns 34%, below 50%, so it still predicts "stayed" for everyone. And all of this accuracy is measured on the same 200 users the model was fitted to, with no held-out set, which flatters it further.
The model scores 92.0% accuracy against a 78.0% always-stayed baseline. What is the most careful reading?
Tests · Verify the three-feature fit gives loss about 0.1950 and accuracy 92.0%, and that adding is_mobile gives loss about 0.1823, accuracy 93.5% and a positive is_mobile weight of about 1.04.
The solution prints, for the three-feature model, a loss of 0.1950 and accuracy 92.0%. With is_mobile added the loss drops to 0.1823, accuracy rises to 93.5% and the new weight is +1.044: Mobile users are more likely to churn even after the other features. That agrees with the 30.9% against 11.1% from Part 1, and it is a gain on the training data only.
#Milestone 7: Plan the Follow-Up
The data is observational, so the next step is an experiment. Suppose the team ships a retention change aimed at Free users and hopes to cut churn from 22% to 17%. How many users does the experiment need? The Power, Sample Size and Multiple Testing lesson gives the recipe: fix the significance level (0.05), the power (0.80) and the effect you want to detect, then solve for n.
With z for alpha = 0.05 of 1.960 and z for 80% power of 0.842, the answer is 985 users per arm, 1,970 in total. As a check, the cell simulates 4,000 experiments with a true drop from 22% to 17% at that size, and the test rejects in 80.1% of them, matching the target power. Our whole StreamBox sample has only 200 users, far too small for a 5 point question. Halving the effect, a drop to 19.5%, needs 4,130 per arm, about four times as many.
Now the correction. If the team watches several metrics (churn, sessions per week, minutes per session) and wants a 5% chance of any false alarm, a Bonferroni correction divides alpha by the number of metrics. With 3 metrics each test uses alpha = 0.0167 and needs 1,314 users per arm. With 5 metrics each uses 0.0100 and needs 1,466.
Drag the effect size and the metric count in this power viz to watch the required n move.
Watching 3 metrics with a Bonferroni correction raises the required sample from 985 to 1,314 per arm. Why?
Tests · Verify 985 per arm at power 0.80 and 1,318 per arm at power 0.90 for 22% to 17%, 4,130 per arm for 22% to 19.5%, and the costs 7,880 and 10,544 dollars under the assumed 4 dollars per enrolled user.
The solution prints 985 per arm at power 0.80 and 1,318 at power 0.90, and 4,130 per arm for the smaller 2.5 point effect. Under an assumed cost of 4 dollars per enrolled user (an assumption, not a StreamBox fact), the two designs cost 7,880 and 10,544 dollars. Asking for 90% power instead of 80% costs 2,664 dollars more for the same effect.
#Milestone 8: The Bayesian View
Frequentist tools answer "how surprising is this under no difference?". The Bayesian view answers the question the product manager actually asked: how likely is it that Free churns more than Pro, and by how much? The Bayesian Inference in Practice lesson gives the machinery. A Beta(1, 1) prior is flat, and the data updates it. Free has 34 churners in 100 users, so its posterior is Beta(1 + 34, 1 + 66) = Beta(35, 67). Pro has 10 in 100, so Beta(11, 91).
To compare two uncertain rates, draw from both posteriors and count, which is the Monte Carlo idea from Inequalities and Monte Carlo.
The posterior means are 0.343 for Free and 0.108 for Pro. Across 200,000 paired draws, Pro came out higher in only 1, so P(Free churn > Pro churn) prints as 1.0000 (the exact value is just under it). The posterior probability that Free exceeds Pro by 10 or more points is 0.9925, and the 95% credible interval for the gap is 0.126 to 0.345. That is close to the frequentist confidence interval from Part 1 (0.130 to 0.350), as it should be with 100 users per group and a flat prior.
The answer is a direct statement about the quantity of interest, which is why many analysts prefer it for decision making. It is still conditional on the model and the prior, and it still says nothing about cause.
The posterior gives P(Free churn > Pro churn) of essentially 1. What can you conclude?
#Milestone 9: Estimation and Information
Under every analysis above sits an estimate. The MLE and Fisher lesson gave one principle, maximum likelihood, and the Information Theory lesson gave a way to score a probabilistic prediction. We use both, and this is where we score the model you built in Milestone 6.
Start with the churn probability. Each user is a coin flip with an unknown probability p, and the likelihood is the product of those flips. The maximum is at the sample proportion, 44 out of 200. For the session length, use a Gaussian on the log minutes: the MLE is the sample mean and the standard deviation with divisor n (not n minus 1) of the logged values. Then score predictors: a baseline that ignores everything and always says "22% chance of churn", and your logistic model.
The MLE of the churn probability is 0.220, and a grid search over p lands on the same value, with a log-likelihood of -105.38. For the log session length the MLE is mu = 3.095 and sigma = 0.639, which implies a median session of exp(3.095) = 22.08 minutes, very close to the sample median of 22.15.
The entropy of churn is 0.7602 bits. The constant 22% predictor has a cross-entropy of exactly 0.7602 bits, because its guess equals the true rate. A 50/50 guess scores 1.0000 bit, so it is worse. Your logistic model scores 0.2813 bits (0.1950 nats), so its features supply 0.4789 bits per user. A model that knows only the plan, with rates of 34% for Free and 10% for Pro, scores 0.6969 bits, a small gain over the baseline.
That gap between 0.7602 and 0.2813 is the same number the training loop drove down in Milestone 6, since cross-entropy is the loss. It is also inflated by leakage: most of the 0.4789 bits come from behaviour columns built from the answer, and the plan-only figure of 0.6969 is the honest, pre-churn one.
The constant 22% predictor scores 0.7602 bits and your model scores 0.2813 bits. What does that tell you?
#Milestone 10 (Optional): Weeks as a Markov Chain
This one is for readers who finished the Markov Chains and MDPs lesson. Instead of asking "did the user ever churn?", model each user week by week as being active or churned, with a transition matrix. StreamBox has no weekly records, so the matrix below is an assumption we choose, not something estimated from the data: an active user churns next week with probability 0.05, and a churned user returns with probability 0.10.
The stationary distribution, the left eigenvector of the matrix for eigenvalue 1, is 0.6667 active and 0.3333 churned. A simulation of one user for 10,000 weeks spent 0.6652 of its weeks active and 0.3348 churned. Raising the matrix to the 50th power from an "active" start gives 0.6668 and 0.3332, so the chain has already forgotten where it started. An active stretch lasts 1 / 0.05 = 20 weeks on average.
In steady state a third of the weeks are spent churned, even though only 5% of active weeks start a churn. That is the point of stationary distributions: they depend on the balance of both flows, not on one of them.
The stationary distribution of the chain is 0.6667 active and 0.3333 churned. What does it describe?
#The Final Analyst Memo
An analysis that nobody reads is a private exercise. The last deliverable is a short memo for a product manager who has two minutes and will not run any code. Write it yourself. Start from your Part 1 memo half, add what you learned in this part, and keep four headings in this order. Copy the blank template below into a document and fill it in. There is no model answer to copy, but there is a rubric and a checklist of expected numbers below so that you can mark it.
What we saw
How sure we are
What to do next
What we cannot conclude
#The Marking Rubric
Score your memo on four criteria, 0 to 3 each, for a total out of 12.
| Criterion | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Numbers with units and intervals | Few or no numbers | Numbers without units or without any interval | Most numbers have units and the key ones (churn rate, gap) have intervals | Every key number has units, and the churn rate, the gap and the sample size all carry intervals or ranges |
| Correct interpretation | Confuses p-values, causation or the baseline | One correct reading, one clear error | Readings are correct, but at least one is vague (for example no mention of the baseline) | p-value, interval, posterior and model accuracy are all read correctly, with the baseline named |
| Costed recommendation | No recommendation | A recommendation with no size or cost | A recommendation with users per arm but no cost or no power | Users per arm, effect size, power, number of metrics and a total cost with its assumption stated |
| Honest limits | None | One generic limit | Synthetic and observational named | Synthetic, observational and unmeasured confounders, leakage in the behaviour columns, and what randomisation would fix |
A total of 10 to 12 is a strong memo, 7 to 9 is solid with one gap to fix, 4 to 6 needs another pass on two criteria, and 0 to 3 means start the memo again from the cell outputs.
#Instructor Grading Checklist
Use this to check a student's numbers. Every value below comes from running the lesson's cells. The accept range allows for the Monte Carlo noise of the lighter cells and for sensible rounding.
| Quantity | Expected value | Accept range |
|---|---|---|
| Churn rate | 44 of 200 = 22.0% | exactly 22.0% |
| Churn rate 95% interval | 16.5% to 28.0% | ends within 1 point |
| Median session | 22.15 minutes | 22.1 to 22.2 |
| Mean session and skew | 26.91 minutes, skew 1.84 | 26.9 to 26.95, skew 1.8 to 1.9 |
| Free and Pro churn | 34.0% and 10.0% | exact |
| Plan gap and interval | 24 points, 13 to 35 | 24 exactly, ends within 1 point |
| z-test | z = 4.10, p about 0.00004 | z 4.0 to 4.2, any p below 0.001 |
| Logistic loss, start and after 500 steps | 0.6931 and 0.1950 | 0.693 and 0.19 to 0.20 |
| Finite-difference gradient check | largest difference 5.5e-11 | below 1e-8 |
| Model accuracy and baseline | 92.0% and 78.0% | 91.5 to 92.5 and exact |
| Model cross-entropy | 0.2813 bits, 0.1950 nats | 0.28 to 0.29 bits |
| Users per arm, 22% to 17% | 985 (1,970 total) | 980 to 990 |
| Users per arm, 3 and 5 metrics | 1,314 and 1,466 | within 10 |
| Simulated power at 985 | 0.801 | 0.78 to 0.82 |
| Posterior means | Free 0.343, Pro 0.108 | within 0.005 |
| P(Free > Pro) and 10 point version | about 1 and 0.9925 | above 0.9999 and 0.99 to 0.995 |
| Credible interval for the gap | 0.126 to 0.345 | ends within 0.01 |
| MLE of churn, log-minutes MLE | 0.220 and mu 3.095, sigma 0.639 | exact to 3 places |
| Entropy, 50/50, plan-only | 0.7602, 1.0000, 0.6969 bits | within 0.001 |
| Stationary distribution (optional) | 0.6667 active, 0.3333 churned | exact to 4 places |
Common wrong answers
- Quoting n as the total (1,970) when the question asks per arm, or the reverse.
- Reporting 0.2813 in nats or 0.1950 in bits. Bits use log base 2, nats use the natural log, and 0.2813 bits is 0.1950 nats.
- Using divisor n minus 1 for the Gaussian MLE of the log minutes, which gives a sigma slightly above 0.639.
- Reading a p-value of 0.00004 as the probability that the plans are the same, or the posterior as a share of causation.
- Calling the 92.0% accuracy a success without naming the 78.0% baseline, the leakage or the fact that it was measured on training data.
- Adding the gradient instead of subtracting it, so the loss rises, or skipping standardisation so training crawls.
- Giving only a single number for the churn rate or the gap, with no interval.
- A recommendation with no cost, or a cost with no stated assumption.
Suggested time per milestone
| Section | Minutes |
|---|---|
| Where Part 1 left off | 3 |
| Milestone 6: model churn by hand, with exercise | 18 |
| Milestone 7: plan the follow-up, with exercise | 9 |
| Milestone 8: the Bayesian view | 6 |
| Milestone 9: estimation and information | 8 |
| Milestone 10 (optional): Markov chain | 5 |
| Final memo and self-marking | 11 |
Without the optional milestone the lesson takes about 55 minutes of work, and with it about 60.
#Key Takeaways
- A logistic regression is X @ w, a sigmoid, a cross-entropy loss and the gradient X transposed times (p minus y) over n, and the gradient matches finite differences to 5.5e-11.
- Gradient descent took the loss from 0.6931 to 0.1950 in 500 steps, giving 92.0% accuracy against a 78.0% always-stayed baseline.
- That strong fit is partly an artefact: the behaviour columns were generated from churn, and accuracy was measured on the training users.
- Detecting 22% to 17% churn needs 985 users per arm, and watching 3 metrics with Bonferroni raises that to 1,314.
- P(Free > Pro) is essentially 1 with a credible interval of 12.6 to 34.5 points, and it is a statement about ordering, not cause.
- The base rate is the bar to beat: 0.7602 bits for the constant 22% predictor, against 0.2813 bits for the model.
- Mark the memo against the rubric: numbers with units and intervals, correct interpretation, a costed recommendation and honest limits.
#Quick Check
In the logistic regression, X is 200 by 4 and w has 4 entries. What does X @ w give, and what is the gradient of the loss?