Containerization & Reproducibility
After this lesson, you will be able to:
- Build production-grade ML Docker images with locked dependencies, CUDA-friendly base images, and minimal layer count
- Distinguish 'works on my machine' (Conda envs, virtual envs) from 'works anywhere' (Docker images with pinned versions)
- Use multi-stage builds + .dockerignore + layer caching to keep ML images small and rebuilds fast
- Pick the right base image (PyTorch official, NVIDIA CUDA, distroless) for your training vs inference workload
Before You Start
#Why Containers for ML
A container packages: OS userspace, system libraries (CUDA, cuDNN, NCCL), Python interpreter, packages, your code, and your model artifacts into a single immutable image. Anywhere that runs the container, the environment is identical.
For ML this matters more than for typical web apps because:
- GPU drivers + CUDA + cuDNN versions are tightly coupled
- Floating-point determinism depends on hardware + library versions
- Models are large (GB-scale) and need predictable mounts
- Distributed training requires identical environments across nodes
#ML-Optimized Dockerfile
# Multi-stage build for ML training
# Layer 1: CUDA base
FROM nvidia/cuda:12.4.0-cudnn-devel-ubuntu22.04 AS base
# Layer 2: System deps (rarely changes — caches well)
RUN apt-get update && apt-get install -y --no-install-recommends \
python3.11 python3-pip git curl wget \
&& rm -rf /var/lib/apt/lists/*
# Layer 3: Python deps (changes when requirements.txt changes)
COPY requirements.txt /tmp/
RUN pip install --no-cache-dir -r /tmp/requirements.txt
# Layer 4: Application code (changes every commit)
WORKDIR /app
COPY . /app
# Layer 5: Entry point
CMD ["python", "train.py"]
Key principles
- Order layers by change frequency — system deps rarely change (top), code changes often (bottom). This caches well.
- Use
--no-cache-dirwith pip to avoid bloating the image. - Pin everything —
torch==2.5.1,cuda:12.4.0, exact versions everywhere. - Use
.dockerignore— exclude__pycache__,.git, large datasets. - Multi-stage builds for inference — separate "training" image (with full toolchain) from "serving" image (minimal runtime).
#Multi-Stage Builds for ML
FROM lines. Each FROM opens a fresh image; you can COPY --from=<stage> to pull artifacts from one stage into another. For ML, the canonical use case is keeping the heavy training toolchain out of the serving image.# ---------- Stage 1: build/train (heavy toolchain) ----------
FROM nvidia/cuda:12.4.0-cudnn-devel-ubuntu22.04 AS build
RUN apt-get update && apt-get install -y gcc g++ make python3.11 python3-pip git
COPY requirements-train.txt /tmp/
RUN pip install --no-cache-dir -r /tmp/requirements-train.txt
COPY . /src
WORKDIR /src
RUN python train.py --output /artifacts/model.pt
# ---------- Stage 2: serve (lean runtime) ----------
FROM nvidia/cuda:12.4.0-cudnn-runtime-ubuntu22.04 AS serve
RUN apt-get update && apt-get install -y python3.11 python3-pip && rm -rf /var/lib/apt/lists/*
COPY requirements-serve.txt /tmp/
RUN pip install --no-cache-dir -r /tmp/requirements-serve.txt
COPY --from=build /artifacts/model.pt /model/model.pt
COPY serve.py /app/serve.py
CMD ["python3", "/app/serve.py"]
The serving image only contains: CUDA runtime (not the dev SDK), Python, the inference-side packages, the trained model artifact, and a serving script. You drop gcc, the training data loaders, Jupyter, profilers, etc. Typical size differences in 2026:
| Image | Approximate uncompressed size |
|---|---|
nvidia/cuda:12.4.0-cudnn-devel-ubuntu22.04 (full training toolchain) | 7-9 GB |
nvidia/cuda:12.4.0-cudnn-runtime-ubuntu22.04 (runtime only) | 2-3 GB |
| Distroless Python + ONNX runtime | 600-900 MB |
The size difference matters for two reasons: pull latency on every node (each cold pod has to download the image once), and attack surface (every binary in the image is potentially exploitable). A 7-GB image is ~30s to pull on a 2 Gbps link; a 600-MB image is ~3 seconds.
#Choosing requirements management: requirements.txt vs poetry vs uv vs pixi
The 2024-2026 landscape for Python dependency management exploded with new tools. Here's the practical map:
| Tool | What it is | Sweet spot for ML | Caveats |
|---|---|---|---|
requirements.txt (+ pip-tools) | Plain list of pinned versions | Universally understood; works in every Docker base | Doesn't track transitive locking by default; pair with pip-compile |
poetry | Lockfile + virtualenv + publishing | Library packaging; project metadata in pyproject.toml | Slower to resolve than uv; less attractive for pure training pipelines |
uv (Astral, 2024) | Rust-based pip-replacement + resolver | Fastest installs in 2026 (~10x pip); first-class lockfile; venv management | Newer; some niche packages still need pip fallback |
pixi (Prefix.dev, 2024) | Conda-compatible package manager built on conda-forge | When you need conda-forge packages (CUDA, FFmpeg, RAPIDS) | Smaller ecosystem than pip/uv for pure Python deps |
conda/mamba | Classic conda environment | Legacy ML stacks that need conda-forge binaries | Slower; heavier base images |
uv + a locked requirements.txt. It's pip-compatible, resolves in milliseconds, and produces fully deterministic installs across machines. Reach for pixi only when you need non-Python system binaries that pip can't provide. Reach for poetry when you're packaging a library, not just deploying a training job.# 2026 idiomatic: uv-based dependency install
RUN pip install --no-cache-dir uv==0.5.0
COPY requirements.txt /tmp/
RUN uv pip install --system --no-cache-dir -r /tmp/requirements.txt
#GPU Container Runtime: nvidia-container-toolkit
docker run cannot access the GPU. NVIDIA's container toolkit (nvidia-container-toolkit, formerly nvidia-docker) injects the host GPU drivers and /dev/nvidia* device nodes into the container at start time.# Install nvidia-container-toolkit on the host (one time):
distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list \
| sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update && sudo apt install -y nvidia-container-toolkit
sudo systemctl restart docker
# Now run a container with GPU access:
docker run --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
docker run --gpus '"device=0,1"' my-train-image python train.py # only GPUs 0 and 1
nvidia.com/gpu: 1 in your pod spec and the scheduler places it on a node with available GPUs.#Kubernetes Operators for ML
Beyond the GPU Operator, several Operators handle ML-specific concerns:
| Operator | What it manages | Use when |
|---|---|---|
| NVIDIA GPU Operator | Drivers, container runtime, DCGM monitoring | Any Kubernetes cluster with NVIDIA GPUs |
| Kubeflow Training Operator | Distributed training jobs (PyTorch, TF, MPI, XGBoost) | Multi-node training without writing pod specs |
| KServe | Model serving with autoscaling, canary, A/B | Production inference on K8s |
| Volcano | Gang scheduling for distributed training | Mixed-priority job queues; pre-emption needed |
| Ray Operator | Ray clusters for distributed Python (RLlib, Train, Serve) | Ray-based workloads (RL, embarrassingly-parallel training) |
PyTorchJob, InferenceService, RayCluster), the Operator reconciles it. You don't write 200 lines of pod spec to spin up 8-node distributed training; you write 30 lines of YAML.#Deterministic CI/CD for ML Training
Reproducibility extends to the build pipeline. The minimum bar for a 2026 ML CI/CD pipeline:
- Pin everything in the Dockerfile — base image SHA, system package versions, Python deps via
uv.lockorrequirements.txt. - Tag images by git SHA —
myorg/train:abc123def, notmyorg/train:latest.latestis poison; it's the tag that means "I don't know which version". - Reproducible layers. Use
BUILDKIT_INLINE_CACHE=1and a remote cache so two CI runs of the same git SHA produce byte-identical images. - Sign images —
cosign signafter build; verify on deploy. Required by SLSA Level 3 (and the EU AI Act for high-risk systems). - Don't bake secrets. Use BuildKit
--mount=type=secretfor pip indices or model registry tokens; neverCOPY .npmrc-style secret baking. - Promote images, don't rebuild them — the image that runs in staging is the same SHA that goes to prod. Rebuilding "the same Dockerfile" produces a different image because base image SHAs, transitive dep versions, and apt mirrors drift hourly.
# Skeleton GitHub Actions workflow
name: build-ml-image
on: { push: { branches: [main] } }
jobs:
build:
runs-on: ubuntu-latest
permissions: { id-token: write, contents: read }
steps:
- uses: actions/checkout@v4
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with: { registry: ghcr.io, username: ${{ github.actor }}, password: ${{ secrets.GITHUB_TOKEN }} }
- name: Build & push
uses: docker/build-push-action@v6
with:
push: true
tags: ghcr.io/${{ github.repository }}/train:${{ github.sha }}
cache-from: type=registry,ref=ghcr.io/${{ github.repository }}/cache:train
cache-to: type=registry,ref=ghcr.io/${{ github.repository }}/cache:train,mode=max
provenance: true
- name: Sign
run: cosign sign --yes ghcr.io/${{ github.repository }}/train:${{ github.sha }}
#Base Image Selection
| Image | Use case |
|---|---|
nvidia/cuda:12.4.0-cudnn-devel | Training (full CUDA toolkit) |
nvidia/cuda:12.4.0-cudnn-runtime | Inference (smaller, runtime-only) |
pytorch/pytorch:2.5.0-cuda12.4-cudnn9-runtime | Pre-baked PyTorch + CUDA |
huggingface/transformers-pytorch-gpu | HuggingFace stack |
gcr.io/distroless/python3-debian12 | Minimal Python (no shell), best security |
python:3.11-slim | CPU-only Python apps |
pytorch/pytorch images for prototyping, build custom from nvidia/cuda for production.#Pinning Strategy
Three levels of pinning:
| Level | Mechanism | Pros / Cons |
|---|---|---|
| Loose | torch>=2.0 in requirements | Easy; reproducibility breaks over time |
| Pinned | torch==2.5.1 exact versions | Reproducible until base image changes |
| Locked | torch==2.5.1 + lock file (uv.lock, poetry.lock, requirements.txt pinned) + base image SHA | Fully reproducible |
uv pip compile or pip-tools to generate locked files; reference base images by SHA digest:FROM nvidia/cuda@sha256:abc123def... AS base
#Hands-On
Tests · Verify the image builds successfully. Verify it runs with --gpus all. Verify .dockerignore excludes data and __pycache__. Use 'docker history' to confirm layers are ordered correctly.
#Compute the Dockerfile Size Difference: CUDA-full vs Slim
Let's quantify the layer-by-layer size of a typical ML Docker image and see exactly where the bytes go. This lets you reason about caching, pull latency, and the cost of multi-stage separation.
The 60-70% size reduction between the training and inference image isn't only about disk: it shows up directly in cold-start latency on every new pod, in attack surface (every binary in the image is a potential CVE), and in deploy time. Once you internalize this, multi-stage builds stop being a "best practice" and become the obvious default.
A team's training image is 14 GB. They deploy via Kubernetes; each new pod takes 4-6 minutes to start because the image must be pulled. Most defensible fix?
Your CI builds an image, tags it `:latest`, and pushes. Staging pulls it, looks great. Hours later production pulls `:latest` and behaves differently. What's the most likely cause?
#Key Takeaways
- Containers solve "works on my machine". Package OS, CUDA, Python, deps, code into one immutable image
- Order layers by change frequency. System deps top, code bottom; caches well
- Pin everything. Base image SHA, exact CUDA version, locked Python deps
- Multi-stage builds keep inference images small. Devel toolchain in training, runtime-only in inference
- Modal / Anyscale / Vertex AI abstract container management — start there for new projects
#Quick Check
Why does Dockerfile layer ORDER matter for ML training images?