End-to-End ML Pipeline
After this lesson, you will be able to:
- Follow data through every stage of a real ML pipeline: from raw data to deployed model to ongoing monitoring
- Spot what can go wrong at each stage and build safeguards against it
- Design pipelines that are repeatable, trackable, and automated, not just notebooks that work once and break forever
Before You Start
#The Iceberg of Machine Learning
This is the foundational lesson for everything that follows in this track. Understanding the full pipeline will make every other topic — MLOps, feature stores, monitoring — click into place. You are building a mental map that will guide your entire ML engineering career.
A production ML pipeline is a directed acyclic graph (DAG) of interconnected stages, each with its own inputs, outputs, validation checks, and failure modes. Understanding the full pipeline is what separates ML engineers from ML hobbyists.
#Stage 1: Problem Definition & Scoping
Before writing a single line of code, answer these questions:
Business framing
- What decision will this model inform?
- What is the cost of a wrong prediction? (False positive vs. false negative)
- What is the baseline? (What happens without ML?)
- What latency and throughput requirements exist?
ML framing
- Is this classification, regression, ranking, generation, or recommendation?
- What is the label? How is it defined? Is it noisy?
- What is the target metric? (Precision@k, NDCG, RMSE, F1, BLEU, human preference?)
- How much labeled data exists? Can we get more?
#Stage 2: Data Ingestion & Validation
Raw data arrives from databases, APIs, event streams, data lakes, third-party vendors, and user uploads. This stage handles:
Data sources
- Batch ingestion: periodic pulls from databases or data warehouses (Snowflake, BigQuery, Redshift)
- Streaming ingestion: real-time events via Kafka, Kinesis, or Pub/Sub
- External APIs: third-party data enrichment
- User-generated: uploads, annotations, feedback
Data validation
- Schema validation: column names, types, and constraints match expectations
- Statistical validation: distributions have not shifted beyond thresholds
- Completeness checks: no missing partitions, expected row counts
- Freshness checks: data is not stale
Tests · Verify that the null check correctly identifies the null user_id row. Verify the range check catches the negative value. Add your own duplicate check.
#Stage 3: Feature Engineering
Raw data is rarely suitable for models. Feature engineering transforms raw signals into informative representations:
Common transformations
- Numerical: standardization, log transforms, binning, polynomial features
- Categorical: one-hot encoding, target encoding, embedding lookup
- Temporal: hour-of-day, day-of-week, time-since-last-event, rolling aggregates
- Text: TF-IDF, word embeddings, tokenization for LLMs
- Aggregation: user-level statistics (mean purchase value, click rate over 7 days)
Tools like Feast, Tecton, and Hopsworks provide:
- Consistent features between training and serving (avoiding training-serving skew)
- Point-in-time correctness (no future data leaking into training features)
- Feature reuse across multiple models and teams
A model trained on features computed offline performs well in evaluation but poorly in production. What is the most likely cause?
Training-serving skew is one of the most insidious bugs in ML systems. If your training pipeline computes a feature using a Pandas window function but your serving pipeline uses a SQL query, subtle differences (handling of nulls, edge cases, time zones) will cause the model to see different feature distributions at serving time than it saw during training. Feature stores exist specifically to solve this problem.
#Stage 4: Training & Experimentation
This is the stage most people think of when they hear "machine learning." But even here, production training is vastly different from notebook experimentation:
Reproducibility requirements
- Pinned library versions (pip freeze, conda lock)
- Fixed random seeds for data splits and initialization
- Version-controlled training code, config, and data
- Logged hyperparameters, metrics, and artifacts
Experiment tracking
- Every training run records: hyperparameters, metrics over time, model artifacts, data version, code commit
- Tools: MLflow, Weights & Biases, Neptune, Comet
- Enables comparison across runs and reproduction of any result
Distributed training patterns
- Data parallelism: replicate model across GPUs, split data
- Model parallelism: split model across GPUs (for models too large for one GPU)
- Pipeline parallelism: different layers on different GPUs, micro-batching
- Frameworks: PyTorch FSDP, DeepSpeed, Horovod
#Stage 5: Evaluation & Validation
A model that performs well on a test set is necessary but not sufficient for production deployment:
Offline evaluation
- Held-out test set performance (the standard)
- Stratified evaluation: performance across demographic groups, edge cases, rare classes
- Regression testing: new model must not degrade on scenarios the old model handled well
- Behavioral testing: specific input-output expectations (like unit tests for models)
Model validation gates
- Performance threshold: accuracy/F1/NDCG above minimum
- Fairness check: performance parity across protected groups
- Latency check: inference time within SLA
- Size check: model fits within serving infrastructure constraints
- Signature check: input/output schema matches serving expectations
Your new model has 2% higher accuracy than the current production model on the test set. Should you deploy it?
A 2% overall accuracy improvement could mask a 15% degradation on your most important customer segment. Always evaluate on slices, not just aggregates. Production model validation is a multi-dimensional decision, not a single number.
#Stage 6: Deployment & Serving
Getting a model from a training artifact to serving production traffic involves:
Deployment strategies
- Shadow mode: New model runs alongside production model, predictions logged but not served. Compare outputs.
- Canary deployment: Route 1-5% of traffic to new model. Monitor metrics. Gradually increase.
- Blue-green: Two identical environments. Switch traffic from blue (old) to green (new) atomically.
- A/B testing: Randomly split users between model versions. Measure business metrics.
Serving patterns
- Online (real-time): REST/gRPC endpoint, sub-100ms latency
- Batch: process millions of records on a schedule, write results to a database
- Streaming: consume events from Kafka, produce predictions in near-real-time
- Edge: model runs on device (mobile, IoT) — ONNX, TFLite, Core ML
#Stage 7: Monitoring & Feedback Loops
Deployment is not the finish line — it is the starting line. Production models degrade over time:
What to monitor
- Data drift: Input distribution shifts (new user demographics, seasonal changes)
- Concept drift: Relationship between inputs and outputs changes (user preferences evolve)
- Prediction drift: Model output distribution shifts
- Performance metrics: Accuracy, latency, throughput, error rates
- Business metrics: Revenue, engagement, customer satisfaction
Drift detection tooling (2024-2026 landscape)
| Tool | Open source? | Strongest feature |
|---|---|---|
| Evidently AI | Yes | Comprehensive drift report library (PSI, KS, Wasserstein, JS) with HTML dashboards |
| NannyML | Yes | Estimated performance without ground-truth labels (CBPE, DLE) — rare and powerful |
| Arize AI / Phoenix | Hybrid | Embedding drift via UMAP, slice-level monitoring |
| WhyLabs | Hybrid | Privacy-preserving profile-based monitoring (no raw data leaves your VPC) |
| Fiddler AI | SaaS | Explainable drift — attributes drift to specific features |
| Soda / Great Expectations | Yes | Data-side checks (schema, null rates, ranges) at ingestion |
#Drift Math: PSI and KS in Practice
Pick a statistic, set a threshold, alert when it breaks. PSI is the workhorse for tabular features; KS is the workhorse for continuous distributions. Run them yourself.
Feedback loops
- Implicit: user clicks, purchases, time-on-page
- Explicit: thumbs up/down, star ratings, corrections
- Human-in-the-loop: flagged predictions reviewed by domain experts
#The Complete Pipeline: A Visual Summary
Here is the end-to-end ML pipeline as a continuous cycle. Each stage feeds into the next, and monitoring feeds back into data ingestion to trigger retraining:
#Step 1: Data Ingestion
Raw data arrives from databases, APIs, event streams, and data lakes. This stage handles batch pulls (Snowflake, BigQuery), streaming ingestion (Kafka, Kinesis), and external enrichment. Every batch goes through schema validation, completeness checks, and freshness verification before proceeding.
#Step 2: Feature Store
Raw data is transformed into informative features — numerical scaling, categorical encoding, temporal aggregations, text embeddings. A feature store (Feast, Tecton) ensures identical feature computation between training and serving, preventing the dreaded training-serving skew.
#Step 3: Training Pipeline
Features are pulled from the store and fed to the model. Every run logs hyperparameters, metrics, code version, and data version to an experiment tracker (MLflow, W&B). Distributed training on GPU clusters with checkpointing handles large-scale jobs.
#Step 4: Model Registry
Trained model artifacts are versioned and stored in a registry. Each version is tagged with its training run, dataset version, and evaluation metrics. The registry enables promotion workflows: Development to Staging to Production, with automated gates at each transition.
#Step 5: Deployment
Models move from the registry to serving infrastructure via shadow mode (run alongside production without serving users), canary deployment (route 1-5% of traffic), or blue-green switching. Each strategy trades speed against safety.
#Step 6: Monitoring
Track data drift (PSI, KS tests), prediction drift, model confidence, latency, throughput, and business metrics. Monitoring detects when the world changes and the model needs retraining. Alerts fire when distribution shifts exceed thresholds.
#Step 7: Retraining Trigger
#Pipeline Anti-Patterns
#Key Takeaways
- The ML code is the tip of the iceberg. Training is a tiny fraction of a production ML system; the majority of effort goes into data pipelines, feature stores, serving infrastructure, monitoring, and testing
- Every pipeline stage has distinct failure modes. Data ingestion can fail silently (schema changes), training can diverge (hyperparameter issues), and serving can degrade (distribution drift); build defenses at each stage
- Reproducibility requires versioning everything. Data, code, model weights, hyperparameters, and environment configurations must all be versioned together to reproduce any past result or debug any production issue
- Automate pipelines from day one. Jupyter notebooks that run once are not pipelines; production systems need orchestrated, scheduled, idempotent workflows that run reliably without human intervention
#Quick Check
What is 'training-serving skew'?