Train/Val/Test & Data Leakage
After this lesson, you will be able to:
- Understand why you need three separate piles of data (training, validation, test) and what each one is for, and why looking at test results to make model decisions contaminates the evaluation
- Split your data correctly for different situations — random splits, stratified splits for class imbalance, grouped splits for entity data, and chronological splits for time series
- Spot and prevent the five types of data leakage — target leakage, train-test contamination, temporal leakage, group leakage, and external data leakage
- Apply sklearn Pipelines to make preprocessing leakage structurally impossible by ensuring scalers and encoders only fit on training data within each cross-validation fold
Before You Start
Don't worry if train/test/validation feels confusing at first — it is really just the idea of practicing on some problems and saving others for the real test. You already do this when you study!
#Taking a Test with the Answer Key
The train/val/test split is not just a best practice — it is the foundation of honest model evaluation. Get it wrong and nothing else matters: your accuracy numbers are fiction, your model comparisons are meaningless, and your production deployment is a time bomb.
#The Three-Way Split
Your model gets 99% accuracy on the test set. Is it good?
When a model achieves suspiciously high accuracy, the first question should always be: "Is there leakage?" Real-world ML problems rarely yield 99% accuracy. If your model seems too good to be true, it almost certainly is.
#Training Set (60-80% of data)
The model learns from this data. It sees these examples repeatedly during training, adjusting its parameters to minimize loss on these examples.
#Validation Set (10-20% of data)
Used to tune hyperparameters and make modeling decisions. The model never trains on this data, but you use it to compare models, select features, and decide when to stop training. Think of it as a practice test — it guides your studying, but it is not the real exam.
#Test Set (10-20% of data)
Try it! Open the Python REPL and type these lines yourself. Split a list into train and test:data = list(range(100)); train, test = data[:80], data[80:]; print(f"Train: {len(train)} examples, Test: {len(test)} examples")— you just did your first train/test split!
#The Splitting Process, Step by Step
Here is exactly how a proper train/validation/test split unfolds, from a single dataset to three non-overlapping subsets used at different stages of model development:
#Step 1: Full Dataset
You begin with a single labeled dataset. Every row has features (inputs) and a target (output). Before doing anything else, you need to decide how this data will be divided. The goal: create three subsets that serve three distinct purposes.
#Step 2: Shuffle Randomly
Randomly shuffle the rows to break any ordering artifacts (alphabetical, chronological, or by data source). Without shuffling, the first 70% might come from one source and the last 30% from another — creating a biased split. For time series data, skip this step and use chronological ordering instead.
#Step 3: Split into Train (70%)
The first and largest portion becomes the training set. The model sees these examples repeatedly during training, adjusting its parameters to minimize loss. This is the data the model learns from.
#Step 4: Split Validation (15%)
The next portion becomes the validation set. Used to tune hyperparameters, compare model architectures, and decide when to stop training. The model never trains on this data, but you use validation results to make modeling decisions.
#Step 5: Hold Out Test (15%)
The final portion is locked away as the test set. This data is used exactly once — after all modeling decisions are finalized. It provides the only unbiased estimate of how your model will perform on truly unseen data.
#Step 6: Train on Train Set
Feed the training set to your model. The model iterates over these examples multiple times (epochs), adjusting weights to minimize the training loss. Monitor both training loss and validation loss during this process.
#Step 7: Tune on Validation
After each training iteration, evaluate on the validation set. Use validation performance to tune learning rate, regularization strength, number of layers, and other hyperparameters. Stop training when validation performance plateaus or degrades (early stopping).
#Step 8: Final Evaluation on Test
Only after your model is finalized, run it on the test set. This single number is your honest performance estimate. If you change anything after seeing test results, the test set is contaminated and you need a new one.
#Why Not Just Two Sets?
Without a validation set, you use the test set to tune hyperparameters. Each time you check test performance and adjust, you implicitly leak test information into your modeling decisions. After dozens of iterations, the test set is no longer an unbiased estimate of real-world performance.
The validation set is your "expendable" evaluation data — you can look at it as many times as you want without compromising the test set's integrity.
#Splitting Strategies
#Random Split
Shuffle the data and divide into proportions. The default for i.i.d. (independent and identically distributed) data.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)#Stratified Split
Ensures each split has the same class distribution as the full dataset. Critical for imbalanced data.
#Grouped Split
When multiple examples come from the same source (same patient, same user, same experiment), they must all go into the same split.
from sklearn.model_selection import GroupShuffleSplit
gss = GroupShuffleSplit(n_splits=1, test_size=0.2)
train_idx, test_idx = next(gss.split(X, y, groups=patient_ids))#Time Series Split
Time -------->
[---- Train ----][-- Val --][-- Test --]
Past Recent Future
Fold 1: [Train ][Val]
Fold 2: [ Train ][Val]
Fold 3: [ Train ][Val]
#Cross-Validation
- Split data into K equal folds (typically 5 or 10)
- For each fold: train on K-1 folds, validate on the remaining fold
- Average the K validation scores
#Data Leakage: The Silent Killer
Data leakage occurs when information from outside the training data is used to create the model. It makes models appear much better during evaluation than they actually are in production.
#Leakage Type 1: Target Leakage
A feature that is a direct consequence of the target variable, not a cause. Example: predicting hospital readmission using "discharge_date" — patients who were readmitted have shorter discharge-to-readmission gaps, but you would not know the readmission date at prediction time.
Another example: predicting whether a loan will default using "number_of_collection_calls." Collections happen AFTER default — this feature does not exist at the time you need to make the lending decision.
#Leakage Type 2: Train-Test Contamination
Test data information leaking into training. Examples:
- Normalizing features using the full dataset (including test data) before splitting
- Computing target encoding statistics on the full dataset
- Imputing missing values using the full dataset's median
- Fitting a PCA on the full dataset before splitting
#Leakage Type 3: Temporal Leakage
Using future information to predict the past. Examples:
- Predicting Monday's stock price using features computed from Tuesday's data
- Predicting patient outcome using lab results from after the outcome occurred
- Training on shuffled time series data (breaks temporal ordering)
#Leakage Type 4: Group Leakage
Data from the same entity appearing in both train and test. Examples:
- Multiple X-rays from the same patient split across train and test
- Multiple transactions from the same credit card in both sets
- Augmented copies of the same image in both sets
The model learns patient/user-specific patterns instead of generalizable ones. Performance on the test set is inflated because the model recognizes "familiar" entities.
#Leakage Type 5: External Data Leakage
Information leaking through external datasets or features. Examples:
- A competition dataset where the row order correlates with the target (row IDs leak information)
- Joining training data with an external dataset that contains the target variable
- Using a feature that was engineered from the target in a previous analysis step
This is the hardest leakage to detect because the information path is indirect.
#Detecting Leakage
How do you know if you have leakage? Look for these red flags:
- Suspiciously high accuracy: If your model achieves 99%+ accuracy on a problem that domain experts say is hard, something is probably wrong.
- A single feature dominates: If one feature has 10x the importance of all others combined, investigate whether it is a proxy for the target.
- Training and test performance are nearly identical: Some gap is normal (model generalizes imperfectly). If there is zero gap, the model might be seeing test data during training.
- Performance drops drastically in production: The classic leakage symptom. The model was evaluating on leaked information that does not exist at prediction time.
#Try It Yourself
Tests · Verify that the 'right way' uses only training statistics for scaling. Show that leakage inflates test accuracy.
#Special Splitting Scenarios
#Highly Imbalanced Data
#Multi-Label Data
When examples have multiple labels, stratify on the rarest label or use iterative stratification (scikit-multilearn).
#Hierarchical Data
Hospital data: patients nested within hospitals. Split at the hospital level, not the patient level, to test generalization to new hospitals.
#Geographical Data
If your model will be deployed in a new city, your test set should contain data from cities not seen during training. Spatial autocorrelation means nearby locations have similar data — random splitting would leak spatial information.
#Time-Series Splitting: Never Shuffle Temporal Data
from sklearn.model_selection import TimeSeriesSplit
# Time series cross-validation: always respects temporal order
tscv = TimeSeriesSplit(n_splits=5)
for train_idx, test_idx in tscv.split(X):
X_train, X_test = X[train_idx], X[test_idx]
y_train, y_test = y[train_idx], y[test_idx]
# Train always comes BEFORE test in timeYour model has 99% accuracy on test data but 60% on real users. What went wrong?
The most likely culprit is data leakage. The 99% accuracy was never real — it was inflated by information that leaked from the test set into training. Once the model faces genuinely unseen data from real users, the illusion shatters. This is why the gap between test accuracy and production accuracy is the most important diagnostic signal in ML.
#Data Leakage in Action: A CodePlayground
Tests · Verify that the 'good' method uses only training statistics. Demonstrate the difference in accuracy between leaked and clean approaches.
#The Leakage Checklist
Before trusting any model evaluation, run through this checklist:
- Was the data split BEFORE any preprocessing (scaling, encoding, imputation)?
- Were preprocessing parameters (mean, std, encoding maps) fitted on training data ONLY?
- For time series: does the training set precede the test set chronologically?
- For grouped data: are all examples from the same group in the same split?
- Does every feature exist at prediction time? (No future features, no target proxies)
- Was the test set used for ANY modeling decisions? (If so, it is contaminated)
- Were augmented/synthetic examples generated AFTER splitting?
- Is the model's accuracy realistic for this problem domain?
#Key Takeaways
- Three-way splits serve distinct purposes. Training data teaches the model, validation data tunes hyperparameters and guides decisions, and the test set provides a single final unbiased evaluation that must not be peeked at during development
- Data leakage is the silent killer of ML models. It makes models appear excellent in evaluation but fail catastrophically in production; preprocessing leakage (scaling before splitting) is the most common and most insidious form
- Match your split strategy to your data type. Use stratified splits for imbalanced classes, grouped splits when multiple examples come from the same entity, and chronological splits for time series to prevent temporal leakage
- Preprocessing parameters must come from training data only. Fitting scalers, imputers, or encoders on the full dataset (including test data) leaks information and inflates performance estimates; always fit on train, then transform test
- Suspiciously high accuracy signals leakage. If your model achieves 99% on a problem domain experts consider hard, investigate for target proxies, future features, or train-test contamination before celebrating
#Quick Check
What is the purpose of the validation set?