MLOps: CI/CD for ML
After this lesson, you will be able to:
- Explain why standard software CI/CD is not enough for ML and what extra automation you need
- Set up experiment tracking (using MLflow or W&B) so you can always reproduce any result
- Build a model registry workflow that safely promotes models from 'experimental' to 'serving real users'
Before You Start
#Software Engineering Had the Same Problem — 20 Years Ago
If you have ever lost track of which version of your model is "the good one," or asked yourself "wait, which hyperparameters produced that result?" — you already understand the pain that MLOps solves. This lesson gives you the tools to never have that feeling again.
#The Three Dimensions of ML
The Three Dimensions of ML
#MLOps Maturity Levels
#Level 0: Manual Process
Level 0: Manual Process
- Training is manual (run a notebook)
- No experiment tracking (which run produced this model?)
- Model is deployed by copying files
- No monitoring (is the model still working?)
- Pain point: "I cannot reproduce last month's results"
#Level 1: ML Pipeline Automation
- Training is automated and triggered (schedule or data arrival)
- Experiment tracking records every run
- Automated evaluation gates
- Model registry stores versioned artifacts
- Pain point: "Changing the pipeline itself requires manual work"
#Level 2: CI/CD for the ML Pipeline
- The pipeline code itself is tested and deployed via CI/CD
- Changes to features, training logic, or serving code go through automated testing
- A/B testing infrastructure for model comparison
- Full observability across all three dimensions
- This is the goal, but few organizations achieve it fully
A team at MLOps Level 0 is frustrated that their model performance degraded over 3 months but nobody noticed. What Level 1 capability would have caught this?
#The ML CI/CD Pipeline, Step by Step
In software engineering, CI/CD automates the path from code change to production. In ML, the pipeline is more complex because data and models change too. Here is the full CI/CD cycle for an ML system:
#Step 1: Code Commit
A developer pushes changes to the feature pipeline, training code, or serving logic. This triggers the CI pipeline, just like traditional software. The change could be a new feature definition, a hyperparameter tweak, or a bug fix in data preprocessing.
#Step 2: CI — Lint + Unit Tests
Automated checks run immediately: code linting, type checking, unit tests for feature pipeline functions, integration tests for data connectors. These catch obvious bugs before any expensive training runs. Fast feedback (under 5 minutes) keeps iteration speed high.
#Step 3: Training Pipeline Runs
If CI passes, the training pipeline kicks off. The orchestrator (Airflow, Kubeflow) runs the full DAG: data validation, feature engineering, model training. Every run logs hyperparameters, metrics, and artifacts to the experiment tracker. GPU time is expensive, so this step only runs when CI confirms code correctness.
#Step 4: Model Evaluation vs Baseline
The newly trained model is evaluated against the current production model on a held-out test set, behavioral tests, and fairness checks. Automated gates enforce minimum thresholds: accuracy above 0.85, latency below 100ms, fairness gap below 0.05. If any gate fails, the pipeline halts.
#Step 5: If Better — Promote to Staging
Models that pass all evaluation gates are registered in the model registry and promoted to the Staging environment. In staging, the model is tested with production-like traffic patterns, load tests, and integration tests against live data sources.
#Step 6: A/B Test in Production
The staging model is deployed as a canary, receiving 5-10% of live production traffic. For 24-72 hours, the system compares the canary model's business metrics (engagement, conversion, revenue) against the control (current production model). Statistical significance is required before proceeding.
#Step 7: Full Rollout
If the A/B test shows statistically significant improvement (or at minimum no degradation), the new model is promoted to 100% of production traffic. The old model is archived but kept available for instant rollback. Monitoring continues to watch for delayed degradation.
#Experiment Tracking: The Lab Notebook of ML
Every training run should record:
| What | Why | Tool |
|---|---|---|
| Hyperparameters | Reproduce the run | MLflow, W&B |
| Metrics over time | Compare runs | MLflow, W&B |
| Code version | Link to exact source | Git SHA, DVC |
| Data version | Know what data was used | DVC, lakeFS |
| Environment | Reproduce the environment | Docker, conda lock |
| Model artifacts | Deploy the exact model | MLflow, S3/GCS |
| System metrics | Diagnose training issues | W&B, GPU monitoring |
Try it! Sign up for a free MLflow or Weights & Biases account (both have free tiers). Run a simple training script and log one metric (like accuracy). Then change a hyperparameter and run it again. Open the dashboard and compare the two runs side by side — this is the "aha moment" where you see why experiment tracking matters.
Tests · Run all 3 experiments and verify the comparison table shows the right values. Implement getBestRun and verify it returns the bert-large run for accuracy.
#Model Registry: Version Control for Models
A model registry is to models what Git is to code — a centralized store of versioned model artifacts with metadata:
Model Registry: sentiment-classifier
| Version | Stage | Accuracy | Created |
|---|---|---|---|
| Version 1 | Archived | 0.82 | 2024-01-15 |
| Version 2 | Production | 0.87 | 2024-02-01 |
| Version 3 | Staging | 0.89 | 2024-02-15 |
| Version 4 | Development | 0.91 | 2024-03-01 |
- Development -> Staging: Offline metrics exceed threshold
- Staging -> Production: A/B test shows improvement on business metrics
- Any -> Archived: Superseded by a better version
Modern model registries (2024-2026 landscape)
| Tool | Strengths | When to pick it |
|---|---|---|
| MLflow Model Registry | Self-hosted, framework-agnostic, mature | Default for self-hosted stacks |
| Hugging Face Hub Enterprise | Massive ecosystem, private repos, inference endpoints | If your team already lives on HF |
| Vertex AI Model Registry | Tight integration with Vertex pipelines + endpoints | All-in on Google Cloud |
| SageMaker Model Registry | IAM, lineage, deploy hooks, multi-account | All-in on AWS |
| BentoML / Bento Cloud | Service-oriented packaging (Bentos), built-in serving | Want model + serving config as one artifact |
| KitOps (ModelKit) / Replicate | OCI-compliant model artifacts, push to any registry | GitOps-first teams |
| W&B Models | Tight loop with experiment tracking | Already on W&B for tracking |
#GitOps for Models
The pattern that won for application deployment (declarative manifests in Git, an operator reconciles cluster state) is now standard for ML:
- ModelKit / KitOps packages model + config + dataset references as an OCI image. The same
docker pushtoolchain works for models. - Flux / ArgoCD with custom CRDs (e.g., KServe
InferenceService) reconciles model deployments from Git manifests. - Promotion-by-PR: changing the
image: my-model:v3line in a manifest and merging the PR is the production deploy. Rollback is agit revert. - Canary + shadow modes are declared in the manifest (traffic split percentages), not configured in a UI.
This gives ML the same audit trail, peer review, and instant rollback that backend services have enjoyed for a decade.
Data lineage from the NLP pretraining-pipelines lesson is the upstream half of every model registry entry — without it, you cannot reproduce a model.#Data Versioning: The Missing Piece
Why it matters
- "Which dataset trained the model that is currently in production?"
- "The model degraded — did the data change?"
- "Can I reproduce last month's training run exactly?"
Tools
- DVC (Data Version Control): Git-like commands for data. Stores metadata in Git, large files in S3/GCS.
- lakeFS: Git-like branching for data lakes. Create a branch of your entire data lake, experiment, merge.
- Delta Lake / Iceberg: Table formats with time travel — query the exact state of a table at any past timestamp.
#CI/CD Pipeline Definition
A complete ML CI/CD pipeline in YAML (conceptual):
# ml-pipeline.yaml (conceptual -- actual syntax depends on your orchestrator)
stages:
data-validation:
trigger: on-data-arrival
steps:
- validate-schema
- check-distribution-drift
- verify-completeness
on-failure: alert-data-team
feature-engineering:
trigger: after-data-validation
steps:
- compute-features
- validate-feature-distributions
- write-to-feature-store
training:
trigger: after-feature-engineering OR on-schedule(weekly)
steps:
- pull-features-from-store
- train-model
- log-to-experiment-tracker
resources:
gpu: a100
timeout: 4h
evaluation:
trigger: after-training
steps:
- evaluate-on-test-set
- run-behavioral-tests
- check-fairness-metrics
- compare-against-baseline
gates:
accuracy: ">= 0.85"
latency_p99: "<= 100ms"
fairness_gap: "<= 0.05"
deployment:
trigger: after-evaluation-passes
strategy: canary
steps:
- register-model
- deploy-canary-5-percent
- monitor-24h
- promote-or-rollback
#Key Takeaways
- MLOps extends CI/CD for the unique challenges of ML. In addition to code changes, ML systems must track data changes, model versions, hyperparameters, and metrics, requiring purpose-built tooling beyond traditional DevOps
- Experiment tracking enables reproducibility. Tools like MLflow and Weights & Biases log every hyperparameter, metric, and artifact for every run, so you can always reproduce or compare past experiments
- Model registries provide safe promotion workflows. Models progress through stages (development, staging, production) with approval gates, automated tests, and rollback capabilities at each transition
- Data and model versioning must work together. A model version is meaningless without the specific data version it was trained on; DVC and similar tools version data alongside code in the same Git workflow
#Quick Check
What are the 'three axes' that change in ML systems, requiring more than traditional CI/CD?