Multi-Model Serving & GPU Infrastructure
After this lesson, you will be able to:
- Architect multi-model serving with shared GPU infrastructure (Triton, vLLM, BentoML) to maximize utilization vs single-model deployment
- Apply GPU sharing techniques (MIG, MPS, virtual GPUs) to run multiple workloads on one physical GPU
- Use spot/preemptible instances + autoscaling to cut serving costs 40-70% for non-critical workloads
- Pick the right serving topology: dedicated per-model, multi-model endpoints, model orchestrator + routing
Before You Start
#The GPU Utilization Problem
Single-model serving wastes GPUs. A typical workload:
- LLM inference is memory-bandwidth-bound, not compute-bound
- Burst traffic patterns leave GPUs idle 80%+ of the time
- Each model gets one GPU minimum, even if it only needs 10% of one
#Serving Topology: Three Architectures
| Topology | What it means | Cost | Isolation |
|---|---|---|---|
| Dedicated per-model | One serving deployment per model; one GPU minimum | Highest | Strongest — each model has its own process and memory |
| Multi-model endpoint | One inference server, many models loaded; model selected per request | 3-7x cheaper | Medium — same process; load/unload at request boundaries |
| Orchestrator + routing | A thin router fronts a heterogeneous fleet of dedicated/MME backends; classifies request -> route | Middle | Configurable — mix dedicated for hot models, MME for the long tail |
For most teams, the right answer is the third: a hot core of high-traffic models on dedicated infra, a long tail of niche models on a shared multi-model endpoint, and a router that classifies the request and dispatches.
inbound request
│
▼
+--------+--------+
| classifier / |
| router |
+--------+--------+
/ | \
▼ ▼ ▼
hot model 1 hot model 2 multi-model endpoint
(dedicated) (dedicated) (50+ niche models on 1 GPU)
200 QPS 180 QPS cold + warm tier
#Multi-Model Serving Architectures
#1. Triton Inference Server (NVIDIA)
The production standard. Supports:
- Dynamic batching: combine requests across users for GPU efficiency
- Model ensembling: chain models in pipelines
- Multi-framework: TensorRT, PyTorch, TensorFlow, ONNX, vLLM
- Concurrent execution: multiple models on one GPU via streams
#2. vLLM (LLM-specific)
For LLM serving specifically:
- Continuous batching: requests of varying lengths share GPU efficiently
- PagedAttention: KV cache as a paging system (24× throughput vs naive)
- Multi-LoRA: serve dozens of fine-tuned LoRA adapters on one base model
# Multi-LoRA serving
from vllm import LLM
llm = LLM(model="meta-llama/Llama-3-70b", enable_lora=True, max_loras=8)
# Now serve 8 different LoRA-fine-tuned variants from ONE 70B base model in memory#Model Swapping at Request Boundaries (2026 frontier)
#3. BentoML / SageMaker MME / Vertex AI Multi-Model Endpoints
Higher-level abstractions: define your models, the platform handles GPU sharing.
#GPU Sharing Mechanisms
| Mechanism | What it does | Use when |
|---|---|---|
| MIG (Multi-Instance GPU, NVIDIA H100/A100) | Hardware-level partitioning into 7 isolated instances | Production isolation needed |
| MPS (Multi-Process Service) | Software-level GPU time-slicing | Trusted workloads, no hard isolation |
| vGPU (NVIDIA virtualization) | Hypervisor-level sharing | Multi-tenant cloud environments |
| CUDA streams | Single-process concurrent kernels | Same model, multiple inference threads |
#Autoscaling Patterns
| Strategy | When |
|---|---|
| Reactive (scale on QPS) | General-purpose |
| Predictive (ML-based forecasting) | Predictable diurnal patterns |
| Scheduled | Known traffic patterns (e.g., business hours) |
| Scale-to-zero | Long idle periods (Modal, Anyscale, Cloud Run) |
For LLM serving specifically: scale-to-zero is risky because cold starts on H100 take 30-90s. Use warm pools or aggressive prediction.
#Cold-Start Mitigation
Cold starts are the single most user-visible failure of cost-optimized serving. The three mitigation tiers, in increasing cost:
| Tier | Technique | Cold start | Cost |
|---|---|---|---|
| 0 | Pre-loaded model in GPU memory (always-warm replica) | 0 ms | Highest — pay for idle GPU |
| 1 | Model on host RAM, page in on first request | 200-800 ms | Low — host RAM is cheap |
| 2 | Model on local NVMe, mmap-load on first request | 1-3 s | Lower — NVMe is cheap, no RAM cost |
| 3 | Model on S3/GCS, download on first request | 30-90 s | Lowest — pay only for S3 storage |
Canary deployments and A/B traffic splitting add their own pre-warming needs: before flipping traffic to the new version, send shadow traffic for 5-10 minutes so the new replicas are warm before they take real load. Triton, KServe, and SageMaker all expose explicit warmup hooks.
#Canary + A/B Traffic Splitting
| Pattern | Use | Risk |
|---|---|---|
| Blue/green | Two identical environments; instant traffic flip | Doubles infra for the cutover window |
| Canary 1% -> 5% -> 25% -> 100% | Progressive rollout; auto-rollback on metric regression | Slower to fully deploy; needs solid metric pipeline |
| Shadow traffic | Send copy of prod traffic to new replica; compare outputs offline | Doubles inference cost during shadow window |
| Multi-armed bandit A/B | Traffic split adapts based on observed quality | Requires real-time quality signal |
Canary is the default for ML. Pre-warm new replicas via shadow traffic for ~10 min, ramp from 1% to 100% over 30-60 minutes with auto-rollback on latency/error regressions.
Your team runs a 70B model on a 4-GPU H100 pod. You add three new LoRA adapters per week. After two months, the pod is at 99% memory and you're seeing OOM kills. Most defensible response?
#Spot / Preemptible Strategy
Spot instances are 50-90% cheaper but can be reclaimed with 30s-2min notice. For ML serving:
- Spot for inference replicas: most replicas on spot, ~10-20% on-demand for stability
- Auto-failover: if spot reclaimed, traffic shifts to on-demand replicas seamlessly
- NEVER spot for training: long-running jobs lose hours of progress on preemption (unless you checkpoint aggressively)
#Multi-Model GPU Sharing Cost Calculator
Cost reasoning for multi-model serving is fiddly because the answer depends on traffic mix, baseline utilization, sharing mechanism, and spot/on-demand strategy. Below is a parametric calculator you can twiddle to see how each lever moves the bill.
ceil(N / 7), and multi-LoRA scales as roughly constant (one H100, dozens of adapters) until you hit the per-GPU memory ceiling. Knowing where each curve sits lets you predict the cost crossover points without rerunning the calculator: above ~5-8 models, multi-LoRA wins decisively; below that, MIG or even dedicated may be simpler.A startup is serving one production model on a dedicated H100, 24/7, with bursty traffic and 8% average GPU utilization. Their CFO wants a 50% cost cut. Lowest-risk single change?
#Hands-On
Tests · Verify the calculator shows multi-model + MIG savings of 80%+ vs dedicated on-demand. Verify spot pricing is ~63% cheaper. Verify multi-LoRA on single H100 is the cheapest at fixed accuracy.
#Key Takeaways
- Multi-model serving is 3-7× cheaper than dedicated. Single biggest cost lever in ML production
- Triton + vLLM are the production standards. Triton for general multi-model, vLLM for LLM-specific
- MIG provides hardware-isolated GPU sharing. 7 partitions per H100, predictable performance
- Spot instances cut 50-90% off serving costs. With mixed on-demand for stability
- vLLM's multi-LoRA serves dozens of fine-tunes from one base model in memory. Game-changing for fine-tune-heavy workloads
#Quick Check
Why does multi-model serving on shared GPUs typically save 3-7× on cost vs dedicated per-model deployment?