Handling Imbalanced Datasets
After this lesson, you will be able to:
- Spot when your dataset has way more of one class than another (like 99% not-fraud and 1% fraud) and measure how bad the imbalance is with a clear imbalance ratio
- Fix imbalanced data using tricks like giving rare examples more weight (class_weight), creating synthetic examples (SMOTE/ADASYN), or removing excess majority examples (undersampling)
- Pick the right success metric for imbalanced data (F1, AUC-ROC, and Precision-Recall AUC) because plain accuracy is a liar when one class dominates
- Combine SMOTE with cleaning techniques (SMOTETomek, SMOTEENN) for better decision boundaries, and know when to reframe extreme imbalance as anomaly detection instead of classification
Before You Start
Don't worry if this feels tricky — class imbalance trips up even experienced engineers. The key insight is simple: when one class vastly outnumbers the other, your model takes the lazy shortcut of always guessing the common class. This lesson teaches you how to fix that!
#Why Imbalance Breaks Standard ML
When you train a standard classifier on imbalanced data, the following sequence happens:
- Training sees mostly majority-class examples — if your dataset is 99% "not fraud," most batches contain almost no fraud examples
- Gradient is dominated by majority class — the model updates primarily to classify "not fraud" correctly, since that is where most of the loss comes from
- Minority class is ignored — the model learns that predicting "not fraud" for everything minimizes the loss function
- Accuracy is misleadingly high — because 99% of predictions are correct by simply predicting the majority class
The result: a high-accuracy model that fails at its actual job.
#Detecting and Measuring Imbalance
Before solving imbalance, you need to measure it:
import pandas as pd
import numpy as np
from collections import Counter
# Check class distribution
print(df['label'].value_counts())
print(df['label'].value_counts(normalize=True)) # as percentages
# Compute imbalance ratio
counts = df['label'].value_counts()
imbalance_ratio = counts.iloc[0] / counts.iloc[1]
print(f"Imbalance ratio: {imbalance_ratio:.1f}:1")
# Counter works for any sequence
print(Counter(y_train))Try it! Open the Python REPL and type these lines yourself. Simulate an imbalanced dataset:from collections import Counter; import numpy as np; labels = np.array([0]*990 + [1]*10); print(Counter(labels)); print(f"Imbalance ratio: {990/10:.0f}:1")— see how extreme 99:1 looks!
Severity guidelines
- Mild imbalance (< 4:1): Class weighting usually sufficient
- Moderate imbalance (4:1 to 20:1): Class weighting or SMOTE
- Severe imbalance (20:1 to 100:1): SMOTE + undersampling combination
- Extreme imbalance (100:1+): Anomaly detection framing may be more appropriate than classification
#The Right Metrics for Imbalanced Data
Standard accuracy is useless. Use these instead:
Your fraud dataset is 1:500 imbalanced (0.2% fraud). You try class_weight='balanced' and SMOTE separately. SMOTE gives substantially better F1. Why might class weighting fail here?
#Fix 1: Class Weighting (Free, Try This First)
Class weighting is the simplest fix: tell the algorithm to pay more attention to minority class errors.
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
# sklearn: class_weight='balanced' automatically computes weights
# inversely proportional to class frequency
model = LogisticRegression(class_weight='balanced')
model.fit(X_train, y_train)
# Manual weights: give 10x penalty to minority class errors
model = RandomForestClassifier(class_weight={0: 1, 1: 10})
# In keras/tensorflow:
# class_weight = {0: 1.0, 1: 500.0}
# model.fit(X, y, class_weight=class_weight)weight_class = n_samples / (n_classes * n_samples_in_class). A class with 100 samples in a dataset of 10,000 with 2 classes gets weight 10000 / (2 * 100) = 50. Its misclassification errors count 50x more in the loss function.- Mild to moderate imbalance (up to ~50:1)
- Limited time / simple pipeline (one parameter to add)
- When the minority class has sufficient examples to learn from (even if underrepresented)
#Fix 2: SMOTE: Synthetic Minority Oversampling
#The SMOTE Algorithm
x_i:- Find its k nearest neighbors among other minority samples (default k=5)
- Randomly choose one neighbor
x_nn - Interpolate between them to create a synthetic sample:
#Using SMOTE with imbalanced-learn
from imblearn.over_sampling import SMOTE, ADASYN
from imblearn.pipeline import Pipeline as ImbPipeline
# Basic SMOTE
smote = SMOTE(
sampling_strategy='auto', # resample minority to match majority
k_neighbors=5,
random_state=42
)
X_resampled, y_resampled = smote.fit_resample(X_train, y_train)
print(f"Before SMOTE: {Counter(y_train)}")
print(f"After SMOTE: {Counter(y_resampled)}")
# SMOTE-NC: handles mixed categorical + numerical features
from imblearn.over_sampling import SMOTENC
# categorical_features: indices of categorical columns
smote_nc = SMOTENC(categorical_features=[0, 2, 5], random_state=42)
X_res, y_res = smote_nc.fit_resample(X_train, y_train)#ADASYN: Adaptive Oversampling
ADASYN (Adaptive Synthetic Sampling) is an improvement over SMOTE that generates more samples in regions where the model is confused:
# ADASYN: generates more samples near the decision boundary
# where classification is hardest
adasyn = ADASYN(
sampling_strategy='auto',
n_neighbors=5,
random_state=42
)
X_res, y_res = adasyn.fit_resample(X_train, y_train)#Fix 3: Undersampling the Majority Class
Instead of creating more minority samples, remove majority samples:
#Random Undersampling
from imblearn.under_sampling import RandomUnderSampler
rus = RandomUnderSampler(
sampling_strategy=0.5, # target ratio: minority/majority = 0.5
random_state=42
)
X_res, y_res = rus.fit_resample(X_train, y_train)#Tomek Links
Tomek links removes majority-class samples that are borderline near minority samples — cleaning the decision boundary without discarding distant majority samples:
from imblearn.under_sampling import TomekLinks
tl = TomekLinks()
X_res, y_res = tl.fit_resample(X_train, y_train)
# Only removes majority samples that are nearest neighbors of minority samples#Cluster Centroids
Replace clusters of majority samples with their centroid — information-preserving compression:
from imblearn.under_sampling import ClusterCentroids
cc = ClusterCentroids(random_state=42)
X_res, y_res = cc.fit_resample(X_train, y_train)#Fix 4: Combining Over- and Undersampling
The most powerful approaches combine SMOTE with cleaning of borderline majority samples:
from imblearn.combine import SMOTETomek, SMOTEENN
# SMOTE + Tomek: oversample minority, then clean borderline majority
smote_tomek = SMOTETomek(random_state=42)
X_res, y_res = smote_tomek.fit_resample(X_train, y_train)
# SMOTE + ENN: more aggressive cleaning than Tomek
smote_enn = SMOTEENN(random_state=42)
X_res, y_res = smote_enn.fit_resample(X_train, y_train)#Putting It All Together: Comparing Strategies
#Key Takeaways
- Accuracy is a liar with imbalanced data. A model predicting the majority class always can have 99.9% accuracy while catching zero minority-class examples; use F1, AUC-ROC, or precision-recall AUC as your primary metrics
- Detect imbalance before modeling. Compute class distribution and imbalance ratio; choose your strategy based on severity (mild: class weights; severe: SMOTE; extreme: anomaly detection framing)
- Class weighting is free and should be your first try.
class_weight='balanced'in sklearn adds a single parameter and significantly improves minority-class recall with no extra compute cost - SMOTE creates synthetic minority samples by interpolating between neighbors. It increases the diversity of minority-class training examples; use it after splitting (never before) and prefer imblearn's Pipeline for cross-validation safety
- Always apply resampling inside the training fold only. SMOTE on the full dataset before splitting leaks test information into training and inflates evaluation metrics
#Quick Check
A credit card fraud model has 99.9% accuracy on a 0.1% fraud dataset, but the business team reports it never catches any fraud. What is the most likely explanation?