Capstone Part 1: Describe, Test and Quantify StreamBox
After this lesson, you will be able to:
- Load one dataset and describe a skewed column with the summaries that suit it
- Choose between a Normal and a lognormal model for session length and defend the choice with numbers
- Check whether a group difference survives once you split by a third variable
- Put a bootstrap interval and a formula interval on the same statistic and explain why they agree
- Test a group difference two ways (a z-test and a permutation test) and read the p-value and the interval together
- Explain why a column generated from the outcome can inflate an apparent relationship (data leakage)
Before You Start
#The Project and the Dataset
StreamBox is an imaginary streaming app. Its data is synthetic: a seeded program generated it, it describes no real company, and no claim below is about a real product. That is useful, because we know how it was made and can check whether our methods recover the truth.
You have 200 users, one row each, with five columns.
- plan: Free or Pro.
- device: Mobile or Desktop.
- churned: 1 if the user left the app, 0 if they stayed.
- session_min: average session length in minutes.
- sessions_per_week: a small count.
You also have a population of 10,000 session lengths that stands in for "everyone", so that we can check a sample against a known answer.
users and population from the first one.#A Warning About This Data: Leakage
session_min and sessions_per_week are drawn after the program already knows whether a user churned. A churned user gets a typical session of about 11 minutes and an average of 1.5 sessions a week. A user who stayed gets about 26 minutes and 4.5 sessions a week.So the line "session length is closely tied to churn" is partly an artefact of how the data was made. We wrote the link in ourselves.
The plan and device columns are different. They were fixed before churn, so a link between them and churn is not leakage. We still will not call it causal, as Milestone 3 explains. Wherever the lesson uses the behaviour columns, we will say so.
#Milestone 1: Load and Describe
users (a list of dictionaries) and population (10,000 numbers), then describes the session column with the tools from Quantiles, Shape and Reading Real Data.session_min the mean is 26.91 minutes and the median is 22.15, so the mean sits 4.76 minutes above the median. The standard deviation is 18.07. The quartiles are Q1 = 14.82 and Q3 = 34.50, which gives an IQR of 19.68 and fences at -14.69 and 64.01. The lower fence is below zero, so nothing can be flagged on the low side. On the high side seven users are flagged: 68.6, 78.4, 86.6, 88.7, 91.4, 92.1 and 119.6. Skewness is 1.84, clearly right-skewed. Sessions per week average 3.66.The last two lines compare churners with stayers. Churned users average 15.29 minutes per session and 1.20 sessions a week, and users who stayed average 30.19 minutes and 4.35 sessions a week. That is a big gap, and it is exactly the leakage warning above: we built it in, so it tells you nothing about a real product.
Do not delete the seven flagged users. They are long-session users, not typing errors, and they are exactly the people a streaming service wants to keep. Report the median and the quartiles, and keep the outliers in.
This is the same box plot logic you met earlier. Drag the extreme-user slider and see which statistics move and which refuse to.
The mean session is 26.91 minutes and the median is 22.15. What is the most sensible reading?
#Milestone 2: Look at the Shape
Numbers describe, pictures reveal. Before choosing a model for session length, look at its histogram and ask which distribution could plausibly have made it. The Normal is the first thing most people reach for, and both From Data to Distributions and The Distribution Zoo warn about that reflex.
This viz draws the same StreamBox column with a histogram and a smooth density, so you can see how the bin width changes the story.
Now compare two fitted models. The cell fits a Normal to the minutes and a lognormal (a Normal on the log of the minutes), both by maximum likelihood, and compares how well each explains the data.
The Normal fit has mean 26.91 and standard deviation 18.02, with a log-likelihood of -862.1. The lognormal has log-mean 3.095 and log-sigma 0.639, with a log-likelihood of -813.2. Both models have two parameters, so the 48.9 point gap is a fair win for the lognormal.
The Normal also says something absurd: it predicts that 6.8% of users have a session shorter than zero minutes. And it gets the tail wrong. In the data 6.0% of users are above 60 minutes. The Normal predicts 3.3%, while the lognormal predicts 5.9%. Taking logs also cures the skew: it falls from 1.84 to -0.14.
Which finding is the strongest evidence against a Normal model for session_min?
Now try a small exercise of your own. Each TODO uses one idea from Milestone 1 or 2.
Tests · Verify the Free median is 20.80 and the Pro median is 24.10, the Free fences are about [-10.65, 53.95], and that seven Free sessions are flagged (60.9, 61.0, 68.6, 78.4, 86.6, 88.7, 91.4).
The solution prints Free median 20.80 and Pro median 24.10. For Free users Q1 = 13.57 and Q3 = 29.73, so the IQR is 16.15 and the fences are -10.65 and 53.95. Seven Free sessions fall above the upper fence: 60.9, 61.0, 68.6, 78.4, 86.6, 88.7 and 91.4. Free users have a smaller typical session than Pro, and a long tail like everyone else.
#Milestone 3: Relationships and a Simpson Check
Now the product manager's question. Does churn depend on plan? On device? And here is the trap from Correlation, Causation and Simpson's Paradox: a group difference can shrink, vanish or even reverse once you split by a third variable. Before you believe a plan effect, you must look inside each device. Plan and device are fixed before churn, so this check does not suffer from the leakage problem above.
Free users churn more than Pro users overall. Free users are also much more likely to be on Mobile. When you compare Free with Pro separately within Mobile and within Desktop, what do you expect?
By plan, churn is 34.0% for Free and 10.0% for Pro. By device it is 30.9% for Mobile and 11.1% for Desktop. The four cells are Free Mobile 28 of 70 (40.0%), Free Desktop 6 of 30 (20.0%), Pro Mobile 6 of 40 (15.0%) and Pro Desktop 4 of 60 (6.7%).
Inside each device the gap survives: Free minus Pro is +25.0 points on Mobile and +13.3 on Desktop, against +24.0 overall. Device is clearly related to churn, and Free users are more often on Mobile (70% of Free users against 40% of Pro users), so device is a confounder candidate. But it does not cancel the plan effect. The honest wording is "Free users churn more, on both devices", and not "the plan causes churn".
Free minus Pro is +25.0 points on Mobile, +13.3 on Desktop and +24.0 overall. What does that pattern show?
#Milestone 4: How Sure Are We
Every number so far is a statistic from one sample, so it carries error. The Sampling, Standard Error and the Bootstrap lesson gave two ways to size it. We will use both and see that they agree.
This viz draws repeated samples from the StreamBox population and bootstraps one of them.
The cell resamples the 200 users with replacement 2,000 times and records the mean session and the churn rate in each resample. We use 2,000 resamples instead of the 10,000 you might see elsewhere to keep the cell quick in the browser. The cost is a little Monte Carlo noise: the interval ends can move by roughly a tenth of a minute from run to run, which does not change any conclusion.
For the mean session the bootstrap gives a 95% interval of 24.54 to 29.42 minutes, with a bootstrap standard error of 1.27. The formula s over the square root of n gives 1.28. For the churn rate the observed 0.220 has a bootstrap interval of 0.165 to 0.280 and a bootstrap standard error of 0.0294. The formula root of p(1-p)/n gives 0.0293, and an interval of 0.163 to 0.277.
The two methods agree to two or three digits, which is what you hope for with 200 independent users. The 10,000-user population has a mean of 26.76 minutes, which sits inside the interval. So "churn is 22%" is better reported as "about 22%, give or take 6 points".
The bootstrap standard error of the mean session is 1.27 minutes. Roughly what 95% interval does that suggest around 26.91?
#Milestone 5: Test the Plan Gap
Free churn is 34 of 100 and Pro churn is 10 of 100, a gap of 24 points. The Statistical Inference lesson asks the formal version of the question: if both plans really had the same churn rate, how surprising would a gap this big be?
The two-proportion z-test pools the two groups to get the rate under that "no difference" assumption, then measures the gap in standard errors. The permutation test needs no formula at all: shuffle the plan labels many times and count how often chance alone produces a gap at least as large.
We shuffle 5,000 times here, again to keep the cell quick. That sets a floor on the p-value the test can report: with 5,000 shuffles the smallest nonzero value is 1 in 5,000, which is 0.0002. A result at that floor means "smaller than about 0.001", not "exactly 0.0002".
The pooled rate is 0.220 and its standard error for a difference is 0.0586. The gap of 0.240 is z = 4.10 standard errors, with a two-sided p-value of about 0.00004. A 95% confidence interval for the gap runs from 0.130 to 0.350. The permutation test shuffled the labels 5,000 times and only 1 shuffle gave a gap at least as large, so its p-value is at the floor of 0.0002. Two routes, the same message.
Read the confidence interval as well as the p-value. The p-value says the gap is unlikely to be luck. The interval says the plausible size runs from about 13 to 35 points, which is a wide range. Pro has only 10 churners, so the normal approximation behind the z-test is on the edge of comfortable, and it is reassuring that the permutation test, which does not rely on it, agrees.
The two-sided p-value for the plan gap is about 0.00004. What does it mean?
One more exercise, this time on the device gap. You will use the z-test and the bootstrap together.
Tests · Verify Mobile 34/110 = 0.309 and Desktop 10/90 = 0.111, diff about 0.198, z about 3.36, p about 0.0008, and a bootstrap 95% interval for the difference of about [0.09, 0.30] that excludes 0.
The solution prints Mobile 34 of 110 = 0.309 and Desktop 10 of 90 = 0.111, a gap of 0.198. The z-test gives z = 3.36 and p = 0.00077. The bootstrap interval for the gap is 0.090 to 0.304, which excludes 0, so the device gap is real too. Mobile is clearly the riskier device.
#Your Analyst Memo, First Half
An analysis that nobody reads is a private exercise. Write the first half of the memo now, for a product manager who has two minutes and will not run any code. There is no model answer to copy. Write it in your own words, with your own numbers from the cells above, then compare each number with what the cells printed.
Copy the template below into a text file or a notebook and fill it in. Two to four sentences under each heading is enough.
What we saw
How sure we are
What it might mean
What we cannot conclude yet
#Key Takeaways
- Describe before you model: StreamBox sessions have mean 26.91, median 22.15 and skew 1.84, so the median and quartiles are the honest summary and the seven flagged users stay in.
- Shape decides the model: the lognormal beats the Normal by 48.9 log-likelihood points and the Normal wrongly predicts 6.8% of sessions below zero minutes.
- Check inside subgroups: Free churns 34.0% against Pro's 10.0%, and the gap holds on Mobile (+25.0 points) and on Desktop (+13.3), so there is no Simpson reversal.
- Every number needs an interval: churn of 22.0% runs from 16.5% to 28.0% by the bootstrap, and the bootstrap and the formula agree on the standard error.
- A test needs two routes and an interval: z = 4.10, a permutation p-value at its floor of 0.0002 and a gap interval of 13 to 35 points all say the gap is real and its size is uncertain.
- Watch for leakage: the session columns were generated from churn, so their strong link to churn is partly built in.
#Quick Check
In the StreamBox data, mean session is 26.91 minutes and median is 22.15. Which summary should lead a report to a product manager?