Model Serving & Inference
After this lesson, you will be able to:
- Compare three ways to send predictions (REST, gRPC, and streaming) and pick the best one for your use case
- Explain how LLM serving systems like vLLM keep GPUs busy using batching, memory management, and speculative decoding
- Design a serving setup that balances speed, volume, cost, and reliability
Before You Start
#From Training to Serving: A Different World
You have made it to one of the most practical lessons in ML engineering. If you have ever wondered how AI products actually work in real life (not just in notebooks), this is where it all clicks. Let's dive in.
#Serving Protocols: REST vs gRPC vs Streaming
#REST (HTTP/JSON)
The most common serving pattern. Simple, universal, works with any client:
REST (HTTP/JSON) Flow
Try it! Open a terminal and usecurlto send a JSON request to a free public prediction API (or a local Flask server running a simple model). Watch the request-response cycle in action. Try sending 5 requests in rapid succession and see how response time changes under load — this is exactly the problem that batching solves.
#gRPC (HTTP/2 + Protocol Buffers)
Binary protocol with schema enforcement. 2-10x faster than REST for serialization:
gRPC (HTTP/2 + Protobuf) Flow
#Server-Sent Events / WebSocket Streaming
Essential for LLMs that generate tokens one at a time:
Server-Sent Events Streaming Flow
A user sends a prompt to a ChatGPT-style service. The full response takes 5 seconds to generate. With streaming, the first token appears in 200ms. Without streaming, the user waits 5 seconds for the full response. Which feels faster to the user?
Select a serving protocol above and step through the inference request lifecycle. Compare REST, gRPC, and streaming side by side. Adjust the concurrency slider to see how throughput and latency change under different loads.
#Traditional Model Serving
For non-LLM models (XGBoost, PyTorch classifiers, embedding models):
Serving frameworks
- TorchServe: PyTorch models. Handles batching, model versioning, multi-model serving.
- TensorFlow Serving: TF models. gRPC + REST. Used extensively at Google.
- Triton Inference Server (NVIDIA): Multi-framework (PyTorch, TF, ONNX, TensorRT). GPU optimized.
- BentoML: Framework-agnostic. Python-native. Easy to package and deploy.
- ONNX Runtime: Convert any framework to ONNX format for optimized cross-platform inference.
Key optimizations
- Dynamic batching: Collect individual requests into batches for GPU efficiency. Wait up to N ms or until batch is full.
- Model compilation: Convert PyTorch models to TorchScript or TensorRT for faster inference.
- Caching: Cache frequent predictions (embedding lookups, common queries).
- Multi-model serving: Load multiple models on one GPU, route requests based on model ID.
#LLM Serving: A Different Beast
Large language models have unique serving challenges:
- Memory bound: A 70B parameter model needs ~140GB in FP16 just for weights (plus KV cache)
- Sequential generation: Each token depends on all previous tokens — hard to parallelize
- Variable-length outputs: Cannot predict how long generation will take
- KV cache management: The key-value cache grows linearly with sequence length
#vLLM: PagedAttention
#Continuous Batching
Traditional batching waits for a batch to complete before starting new requests. Continuous batching (iteration-level scheduling) is smarter:
Continuous Batching Timeline
| Step | Slot 1 | Slot 2 | Slot 3 |
|---|---|---|---|
| Step 1 | Req A token 5 | Req B token 3 | Req C token 8 |
| Step 2 | Req A token 6 | Req B token 4 | Req C DONE -> Req D token 1 |
| Step 3 | Req A token 7 | Req B token 5 | Req D token 2 |
When request C finishes, request D immediately takes its slot in the batch. No request waits for the slowest request to finish. This keeps the GPU saturated at all times.
#Speculative Decoding
Use a small "draft" model to generate N candidate tokens quickly, then verify all N in parallel with the large model:
Draft model (fast): generates 5 tokens speculatively
Large model (slow): verifies all 5 in one forward pass
If first 3 are correct: accept 3 tokens in one step (3x speedup!)
If all 5 are correct: accept all 5 (5x speedup!)
This exploits the fact that many tokens are predictable (common phrases, code patterns, structured output). The speedup is proportional to the acceptance rate.
#LLM Serving Systems Comparison
| System | Key Feature | Best For |
|---|---|---|
| vLLM | PagedAttention, continuous batching, multi-LoRA | General LLM serving, highest throughput |
| TGI (Text Generation Inference) | HuggingFace integration, tensor parallelism | HuggingFace models, quick deployment |
| Triton + TensorRT-LLM | NVIDIA optimization, multi-GPU, in-flight batching | Maximum single-request latency optimization |
| Ollama | Local models, easy setup | Development, single-user desktop |
| llama.cpp | CPU inference, quantization (Q4/Q5/Q8) | Edge deployment, no GPU available |
| SGLang | RadixAttention, structured generation, JSON-mode | Multi-turn chat, constrained output |
| TensorRT-LLM | Kernel fusion, INT4 AWQ, FP8 on H100 | Lowest single-request latency on NVIDIA |
| LMDeploy / TurboMind | INT4-AWQ quantization, persistent batching | Throughput-per-dollar on consumer GPUs |
#Inference Disaggregation (2024-2026)
Production systems implementing this pattern:
- Mooncake (Moonshot AI / Kimi) — separates prefill and decode clusters, shares a global KV-cache pool over RDMA.
- DistServe (Peking University, 2024) — disaggregated serving with per-stage goodput optimization, 4-7x more requests per GPU at the same SLO.
- Splitwise (Microsoft) — phase splitting on different hardware classes (compute-heavy prefill on A100, memory-heavy decode on cheaper H100-MIG slices).
- SGLang v0.3+. Adds disaggregated mode plus chunked prefill that interleaves long-prompt processing with active decode batches.
The win is that prefill and decode have wildly different optimal batch sizes and KV-cache footprints. Co-locating them caps throughput at the lower of the two.
A vLLM replica receives a mix of short chat turns (50-token prompts) and long document-summarization requests (8K-token prompts). Without disaggregation, what is the most common symptom users will notice?
#Multi-LoRA and Multi-Tenant GPU Serving
- vLLM
--enable-lora— supports hundreds of concurrent LoRA adapters with negligible overhead. - S-LoRA (Stanford, 2023) — unified paging for adapters and KV cache; 4x more adapters per GPU than naive swapping.
- Punica. Kernel-level fusion of multiple LoRAs in a single CUDA call (SGMV — Segmented Gather Matrix-Vector).
- TGI multi-LoRA. Same pattern for HuggingFace stacks.
This pattern is what makes per-customer fine-tuning economically feasible: 500 enterprise customers can each have a custom adapter served from one shared base-model fleet.
Inference efficiency techniques from the NLP track (KV-cache management, paged attention, speculative decoding) are the foundation for every production serving system covered here. Tensor and pipeline parallelism, originally developed for training, are now standard in inference for 70B+ models.#The Inference Request Lifecycle
Every model prediction follows the same path from request to response. Understanding each step reveals where latency hides and where optimization pays off:
#Step 1: Model Artifact
The journey starts with a trained model artifact stored in a model registry — a file containing the architecture definition, learned weights, and any preprocessing artifacts (tokenizers, scalers). This artifact was produced by the training pipeline and promoted through evaluation gates.
#Step 2: Load into Serving Framework
At startup, the serving framework (vLLM, TorchServe, Triton) loads the model weights into GPU memory. For a 70B model in FP16, this means transferring 140 GB from disk to GPU — a cold start that takes 30-60 seconds. Replicas stay warm to avoid this latency on user requests.
#Step 3: REST/gRPC Endpoint
The model is exposed as a network endpoint. REST (HTTP/JSON) for universality, gRPC (HTTP/2 + Protobuf) for low-latency microservices, or SSE/WebSocket for streaming LLM tokens. A load balancer distributes incoming requests across model replicas.
#Step 4: Request Arrives
A user request hits the API gateway, which handles authentication, rate limiting, and routing. The request is validated (correct schema, input size limits) and placed in a queue. For dynamic batching, the request waits up to N milliseconds for other requests to form a batch.
#Step 5: Preprocess Input
Raw input is transformed into model-ready tensors. For text: tokenization (splitting into subword tokens, converting to IDs). For images: resizing, normalization. For tabular data: feature lookup from the feature store, scaling, encoding. Preprocessing must match training exactly.
#Step 6: Model Inference
The preprocessed tensor batch is fed through the model. For LLMs, this involves iterative autoregressive generation — each token depends on all previous tokens. vLLM uses PagedAttention for memory efficiency and continuous batching to keep the GPU saturated.
#Step 7: Postprocess Output
Raw model output (logits, token IDs, or embeddings) is converted to a human-readable response. For LLMs: decode token IDs back to text. For classifiers: map class indices to labels, attach confidence scores. Safety filters and PII detection run here too.
#Step 8: Return Response
The formatted response is sent back to the client. For streaming, tokens are sent as they are generated via Server-Sent Events. Latency, token counts, model version, and confidence scores are logged for monitoring. The full trace enables debugging any future issues.
#Batching Strategies Deep Dive
Tests · Compare no-batching, static, and dynamic strategies. Verify that dynamic batching provides better throughput than no batching while keeping latency lower than large static batches.
#Inference Optimization Techniques
#Model Compilation
Convert dynamic computation graphs to static, optimized ones:
- TorchScript / torch.compile: PyTorch graph optimization, kernel fusion
- ONNX Runtime: Cross-platform optimized inference
- TensorRT: NVIDIA GPU-specific optimization. Kernel fusion, precision calibration.
- OpenVINO: Intel CPU/GPU optimization
#Quantization at Serving Time
Reduce precision for faster inference (covered in depth in the next lesson):
- FP32 -> FP16: 2x memory reduction, ~2x speedup, minimal quality loss
- FP16 -> INT8: 2x further reduction, requires calibration
- INT8 -> INT4: Another 2x, primarily for LLMs (GPTQ, AWQ)
#KV Cache Optimization
For LLMs, the KV cache is often the memory bottleneck:
- Multi-Query Attention (MQA): Share K/V heads across attention heads (Llama 2 uses GQA, a middle ground)
- Sliding window attention: Only cache the last N tokens (Mistral)
- KV cache quantization: Quantize cached keys/values to INT8 or FP8
- Prefix caching: Hash and reuse the KV cache for common prompt prefixes (system prompts, RAG templates) — vLLM and SGLang ship this out of the box; cache hit rates of 30-60% are typical in chat workloads.
- Cross-request offload: Page cold KV-cache blocks to host memory or NVMe (LMCache, Mooncake) so popular conversations stay resident.
#Sizing Your Serving Fleet: Run the Math
Before signing off on hardware, work the numbers. The cell below estimates KV-cache memory per request, total memory headroom, and the maximum concurrent batch size for a vLLM-style replica — the exact math your platform team needs in their capacity plan.
#Key Takeaways
- Training and serving are fundamentally different workloads. Training optimizes for throughput (process data as fast as possible), while serving optimizes for latency (respond to each request in milliseconds)
- LLM serving uses specialized techniques for efficiency. PagedAttention manages KV cache memory dynamically, continuous batching maximizes GPU utilization, and speculative decoding predicts future tokens to reduce latency
- Choose your serving pattern based on requirements. REST APIs for simple request-response, gRPC for low-latency microservices, and streaming (SSE/WebSocket) for LLM token-by-token generation
- Cost, latency, throughput, and availability form a quadrilateral tradeoff. You cannot optimize all four simultaneously; production systems must make explicit choices about which dimensions matter most for their use case
#Quick Check
What is PagedAttention (used in vLLM) and why does it matter?
#Exit Ticket
Before moving on, you should be able to answer these in your own words. If any feel shaky, scroll back.