Hyperparameter Tuning
After this lesson, you will be able to:
- Pick the right search strategy for your problem (grid search, random search, Bayesian optimization, or Hyperband) based on the size of the search space and your compute budget
- Define search spaces correctly using log-uniform vs uniform distributions, and avoid the most common silent mistake (uniform-tuning a learning rate)
- Use Optuna with TPE to tune real models and prune unpromising trials early, getting better results in a fraction of the compute of grid search
- Tune inside a cross-validation loop without leaking information into the test set, and know when to use nested CV for unbiased reporting
Before You Start
optuna.create_study() and three lines of code, and pick up the theory once you have the muscle memory.#War Story: The Hour That Replaced a Month of "Trying Things"
num_leaves, different learning_rate, different min_data_in_leaf. Each round was an engineer typing values into Slack and waiting for results.import optuna
def objective(trial):
params = {
'learning_rate': trial.suggest_float('lr', 1e-3, 0.3, log=True),
'num_leaves': trial.suggest_int('leaves', 16, 256),
'min_data_in_leaf': trial.suggest_int('min_data', 5, 100),
'feature_fraction': trial.suggest_float('ff', 0.5, 1.0),
'reg_alpha': trial.suggest_float('alpha', 1e-8, 1.0, log=True),
}
model = LGBMRegressor(**params, n_estimators=5000)
return cross_val_score(model, X, y, cv=5,
scoring='neg_root_mean_squared_error').mean()
study = optuna.create_study(direction='maximize')
study.optimize(objective, n_trials=100)learning_rate of 0.012 (the team's grid had been [0.01, 0.05, 0.1], and they had read on a blog that 0.05 was "the best default"). The blog was right on average, wrong for this dataset.Stop blindly using sklearn defaults. The 1-hour workflow below gets you to ~90% of expert performance on almost any tabular problem.
#The Real Job of an ML Engineer
learning_rate, max_depth, n_estimators, C, gamma. The model's parameters (weights, splits, support vectors) are what training learns. The hyperparameters are what you have to learn separately, by trying combinations and watching the validation score move.The naive approach — try a few values you read on Stack Overflow — is what most people do. It is also why most people leave 10–30% of their model's potential performance on the floor.
#The Four Strategies
#1. Grid Search: The One You Should Almost Never Use
Enumerate every combination of values from a fixed list per hyperparameter and evaluate all of them.
#2. Random Search: The Default Baseline
Sample each hyperparameter independently from a distribution and run a fixed number of trials.
Bergstra and Bengio's central observation: in real ML problems, only a handful of hyperparameters drive most of the variance. Grid search wastes the rest of its budget on the irrelevant dimensions. Random search spends those trials on more values of the dimensions that matter.
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import loguniform, randint
space = {
"learning_rate": loguniform(1e-3, 1e-0), # log-uniform: critical for LR
"max_depth": randint(3, 12),
"subsample": loguniform(0.5, 1.0),
}
search = RandomizedSearchCV(model, space, n_iter=100, cv=5)
search.fit(X_train, y_train)The interactive below makes the grid-vs-random difference visceral. Same model, same 2D hyperparameter space — flip strategies and watch how the dots fall on the validation-score landscape. The convergence strip on the right tracks how fast the best score is found.
In the visualization above, you set the budget to 16 trials. Grid search uses a 4x4 lattice and lands on the optimum. Why does random search often beat grid at the same budget on real problems even though grid 'covers' the space more uniformly?
#A concrete 4-grid-point example
learning_rate ∈ {0.01, 0.1} and max_depth ∈ {4, 8}. Four cross-validated trials:| Trial | LR | Depth | CV AUC |
|---|---|---|---|
| 1 | 0.01 | 4 | 0.814 |
| 2 | 0.01 | 8 | 0.842 |
| 3 | 0.10 | 4 | 0.823 |
| 4 | 0.10 | 8 | 0.798 |
(LR=0.03, depth=6) — somewhere between your grid points. Grid search can never see it. Random search with the same 4 trials samples 4 unique LRs and 4 unique depths and has a much higher chance of stumbling close to the optimum. Bayesian search uses results from trial 2 ("depth=8 helped, LR=0.01 helped") to bias trial 5 toward similar regions, and converges faster still.#The 1-Hour Tuning Workflow (the actually-useful checklist)
This is the workflow that gets you 90% of expert performance on almost any tabular problem:
- Split data into train / test, lock the test set away.
- Pick a baseline model. If unsure, gradient boosting (LightGBM or XGBoost) is almost always the right starting point for tabular.
- Wrap preprocessing in a
Pipelineso it fits per-fold (scaler, encoder, imputer all live in the pipeline, never on the full training set). - Define a log-uniform search space for rate-like hyperparameters (LR, alpha, gamma, dropout). Use
trial.suggest_float(..., log=True)in Optuna. - Run 50-100 Optuna trials with
MedianPrunerand 5-fold CV. Set a 60-minute wall-clock timeout. - Evaluate on the test set once. Report this single number as your honest generalization estimate.
That's it. Six steps, one hour, results that beat hand-tuning by 10-30% on typical problems.
#3. Bayesian Optimization: The One You Should Actually Use
The two dominant Bayesian methods in practice:
- Gaussian Process (GP): model the score as a function with a smooth GP prior. Best for small (≤ 20) continuous spaces. Used by
scikit-optimize,BoTorch. - Tree-structured Parzen Estimator (TPE): model the good trials and bad trials as two separate density estimates and sample from where the good density is high relative to the bad density. Handles categorical and conditional hyperparameters cleanly. Used by
Optuna,Hyperopt.
You're tuning 7 hyperparameters on a model where each CV fit takes 3 minutes. You can afford a wall-clock budget of about 4 hours. Which strategy gives the best expected validation score?
#4. Hyperband and Successive Halving: Compute-Aware Tuning
n_estimators, multi-fold CV). For sub-minute trials, classical Bayesian without pruning is usually simpler.You have a 7-dimensional search space (5 continuous, 2 categorical), a budget of 100 trials, and each trial takes 2 minutes. Which strategy gives you the best expected validation score?
#How to Define Search Spaces (the silent killer)
This is where most tuning attempts die quietly.
#The hyperparameter cheat sheet (what actually matters per algorithm)
| Algorithm | Hyperparameters that matter | Distribution |
|---|---|---|
| Linear / Logistic regression | C (or alpha), l1_ratio | log-uniform |
| KNN | n_neighbors, weights, metric | int / cat |
| SVM (RBF) | C, gamma | log-uniform |
| Decision tree | max_depth, min_samples_leaf, min_samples_split | int |
| Random forest | n_estimators, max_depth, max_features, min_samples_leaf | int / fractional |
| Gradient boosting (XGB / LGBM / CB) | learning_rate, n_estimators (with early stopping), max_depth / num_leaves, min_child_weight / min_data_in_leaf, reg_alpha, reg_lambda, subsample, colsample_bytree | log-uniform / int |
n_estimators should usually be set very large (e.g. 5000) and pinned by early stopping on a validation fold rather than tuned directly. Tune learning_rate instead and let early stopping pick the right number of trees for that LR.#Tuning Inside Cross-Validation
Interactive Lab
Step through K-Fold, Stratified, and TimeSeriesSplit fold-by-fold. The wrong splitter inside a tuning loop is the most common silent leak in production ML — this animation makes the difference obvious.
Pipeline:from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import loguniform
pipe = Pipeline([
("scaler", StandardScaler()),
("clf", LogisticRegression(max_iter=2000)),
])
space = {"clf__C": loguniform(1e-4, 1e2), "clf__l1_ratio": [0.0, 0.5, 1.0]}
search = RandomizedSearchCV(pipe, space, n_iter=80, cv=5, scoring="roc_auc")
search.fit(X_train, y_train)StandardScaler is fit inside each CV fold, on training data only. No leakage.#When you need nested CV
from sklearn.model_selection import cross_val_score, KFold
outer = KFold(n_splits=5, shuffle=True, random_state=0)
unbiased_scores = cross_val_score(search, X_train, y_train, cv=outer, scoring="roc_auc")
print(f"Nested CV: {unbiased_scores.mean():.3f} ± {unbiased_scores.std():.3f}")You are tuning a fraud-detection model on transactions from Jan-Aug, and you'll deploy in Sep. Which CV splitter should you wrap inside your Optuna objective?
Multiple measurements come from the same patient. You're tuning a clinical risk score with stratified k-fold. What's likely going wrong?
#Putting It Together
Tests · Verify random search returns a valid roc_auc score. Verify the test-set evaluation runs after tuning, not during.
#Pruning, Warm Starts, and Multi-Fidelity
A few production-grade tricks that compound:
- Pruning: With Optuna's
MedianPrunerorHyperbandPruner, a trial reports its intermediate score (e.g. score after 10/50/100 trees) and is killed if it falls below the median of past trials at the same step. Typical compute savings: 30–60%. - Warm starting: Save your
studyto a database (storage="sqlite:///study.db") and resume next week with prior trials informing TPE. This compounds across projects. - Multi-fidelity: Tune on a 10% subsample first to find a good neighborhood, then do a small refinement run on the full data with the narrowed search space.
A reviewer asks 'what's your model's expected AUC on truly unseen data after the full tuning procedure?' You ran 100 Optuna trials with 5-fold CV and report the best CV AUC. Why is that number wrong?
You set 'n_estimators': trial.suggest_int('n_est', 100, 2000) AND 'learning_rate': trial.suggest_float('lr', 0.001, 0.5, log=True) in the same Optuna study for XGBoost. What's the more idiomatic choice?
#Key Takeaways
- Random search beats grid search at the same budget. Bergstra & Bengio's 2012 paper proved this for any realistic problem; the only reason to grid-search is small reproducible sweeps for papers
- Bayesian optimization (TPE via Optuna) is the practical default. It uses every past trial to inform the next, typically delivering 5–15% lift over random at the same compute on multi-dimensional spaces
- Use log-uniform distributions for rate-like hyperparameters. Learning rate, C, alpha, gamma, dropout — sampling uniformly is the most common silent bug in tuning, often costing more than picking the wrong algorithm
- Tune inside cross-validation, never on the test set. Wrap preprocessing in a
Pipelineso it fits per-fold; reach for nested CV when the tuning procedure's score itself needs to be reported unbiased - Pruning compounds with Bayesian search. Hyperband / median pruning kills underperforming trials early, often 30–60% compute saving with no quality loss; combine with TPE for the strongest off-the-shelf setup
#Quick Check
Why does random search outperform grid search in practice for the same number of trials?