System Design Case Studies
After this lesson, you will be able to:
- Break down a big AI system design problem into clear steps: requirements, architecture, components, and trade-offs
- Design the full ML architecture for a recommendation system, a ChatGPT-scale chat app, and an autonomous coding agent
- Weigh real engineering trade-offs: speed vs volume, accuracy vs cost, complexity vs maintainability
Before You Start
#How to Approach AI System Design
This is the capstone lesson of the track. Everything you have learned — pipelines, serving, monitoring, cost, security — comes together here. If system design feels intimidating, remember: it is just applying the skills you already have in a structured way.
#Case Study 1: Design Netflix-Style Recommendations
#Step 1: Requirements
Functional
- Recommend movies/shows on the homepage (personalized rows: "Because you watched X", "Top picks for you", "Trending")
- Recommend "More like this" on a title's detail page
- Rank search results by relevance + personalization
- Re-rank continuously (new viewing data, time of day, device)
Non-functional
- 200M+ users, 15,000+ titles
- Homepage must load in under 1 second (recommendations pre-computed or cached)
- Serve ~1B recommendation requests per day
- Support A/B testing (hundreds of concurrent experiments)
- Cold start: handle new users and new titles with no history
Business metrics
- Primary: hours of engagement (not just clicks)
- Secondary: retention rate, trial-to-subscription conversion
- Guard rails: content diversity (avoid filter bubbles), new title discovery rate
Netflix needs to recommend from 15,000 titles to 200M users. Why can't you just score every title for every user in real-time?
#Step 2: Architecture
#Step 3: ML Components
Candidate generation (broad retrieval)
- Collaborative filtering: matrix factorization (user-item embeddings)
- Content-based: item feature embeddings (genre, director, cast, plot)
- Popularity-based: trending titles, regional favorites
- Sequence model: predict next watch from viewing history (Transformer-based)
- ANN (Approximate Nearest Neighbor): retrieve top-500 candidates per user
Ranking model (precise scoring)
- Deep neural network with hundreds of features:
- User features: viewing history, demographics, device, time of day
- Item features: genre, age, popularity, freshness
- Cross features: user-genre affinity, time-since-release decay
- Context features: day of week, holiday, device type
- Trained on: (user, item, context) tuples mapped to engagement scores
- Loss: combination of click probability and expected watch time
Business rules layer
- Diversity: limit titles from same genre in a row
- Freshness: boost newly added content
- Contractual: promote content with expiring licenses
- Parental controls: filter based on user profile
- Exploration: 10% of slots reserved for novel recommendations (explore-exploit)
#Step 4: Trade-offs
| Decision | Option A | Option B | Netflix's Likely Choice |
|---|---|---|---|
| Recompute frequency | Real-time scoring | Batch pre-computation | Hybrid: batch candidates, real-time ranking |
| Model complexity | Simple (logistic regression) | Complex (deep network) | Deep network (worth the complexity at scale) |
| Cold start | Show popular items | Ask preferences | Both: popular items + onboarding quiz |
| Latency vs accuracy | Simpler model, faster | Better model, slower | Pre-compute to get both |
#Case Study 2: Design a ChatGPT Clone
#Step 1: Requirements
Functional
- Multi-turn conversational AI with memory within a session
- Support text generation, code generation, reasoning, summarization, translation
- Stream tokens to user in real-time
- Support system prompts and custom instructions
- Handle file uploads (PDF, images)
- Provide web search and tool use
Non-functional
- 100M+ users, 10M+ daily active
- Target: 100K+ concurrent sessions
- Time-to-first-token (TTFT): under 500ms P95
- Generation speed: 30+ tokens/second
- 99.9% availability
- Cost: under $0.01 per average conversation
Safety requirements
- No harmful content generation
- No PII leakage from training data
- Robust against prompt injection and jailbreaks
- Content filtering for child safety
#Step 2: Architecture
#Step 3: ML Components
The LLM cluster
- Primary model: fine-tuned LLM (70B-400B parameters)
- Alignment: RLHF/DPO trained for helpfulness, harmlessness, honesty
- Serving: vLLM or TensorRT-LLM on H100 clusters
- Multi-GPU: tensor parallelism (2-8 GPUs per model replica)
- Continuous batching for maximum throughput
Session management
- Conversation history stored in Redis (fast access) + PostgreSQL (durable)
- Context window management: when conversation exceeds model context, summarize older turns
- System prompt + user custom instructions prepended to every request
Tool use system
- Function calling: model decides when to invoke tools
- Web search: query → search API → snippet retrieval → inject into context
- Code execution: sandboxed Docker container, timeout limits
- File processing: PDF parsing, image analysis (multimodal model or OCR)
Safety system
- Input classifier: detect jailbreak attempts, harmful requests
- Output classifier: detect harmful, biased, or inappropriate generations
- PII detector: scan outputs for personal information
- Moderation API: complementary rule-based + ML safety checks
#Step 4: Trade-offs
| Decision | Option A | Option B | Recommended |
|---|---|---|---|
| Model size | Smaller (cheaper, faster) | Larger (smarter) | Multi-model routing by difficulty |
| Context handling | Truncate old messages | Summarize old messages | Summarization (preserves context) |
| Safety filtering | Pre-generation (block input) | Post-generation (filter output) | Both (defense in depth) |
| Streaming | Full response at once | Token-by-token | Streaming (better UX) |
| Tool execution | Parallel (fast) | Sequential (safe) | Parallel with dependency analysis |
#Case Study 3: Design an Autonomous Coding Agent
#Step 1: Requirements
Functional
- Accept a natural language task description (e.g., "Add pagination to the /users endpoint")
- Understand the existing codebase (read files, search code)
- Plan and execute multi-step changes across multiple files
- Run tests, read output, fix failures
- Commit changes with meaningful commit messages
- Handle ambiguity by asking clarifying questions
Non-functional
- Support codebases up to 100K files
- Complete tasks in under 10 minutes for most changes
- 80%+ success rate on well-defined tasks
- Cost: under $1 per average task
- Operate safely (never delete production data, never push to main without review)
Safety requirements
- Sandboxed execution (cannot access network, only the repository)
- Human approval for destructive operations (git push, database changes)
- Rate limits on file modifications
- Audit trail of all actions taken
#Step 2: Architecture
#Step 3: ML Components
The reasoning engine
- Core LLM: frontier model with strong coding ability (Claude 3.5 Sonnet, GPT-4o)
- Long context: 100K+ tokens to hold codebase context
- Structured output: tool calls must follow defined schemas
- Multi-turn: maintain conversation state across the entire task
Codebase understanding
- Indexing: embed all files in a vector database for semantic search
- AST parsing: understand code structure, find function definitions, trace call graphs
- Symbol resolution: map "the users controller" to the actual file path
- Context selection: intelligently choose which files to read (not the entire codebase)
Planning and execution
- Task decomposition: break "add pagination" into sub-tasks
- Dependency analysis: determine execution order (modify model → update controller → update tests)
- Error recovery: when tests fail, analyze error, modify code, retry (up to N attempts)
- Verification: run the full test suite after changes, ensure no regressions
Tool definitions
file_read(path): Read file contentsfile_write(path, content): Write file (with diff preview)file_search(query): Semantic search across the codebasegrep(pattern, path): Text search with regexbash(command): Execute shell command in sandboxtest_run(path?): Run test suite or specific test filegit_commit(message): Stage changes and commit
#Step 4: Trade-offs
| Decision | Option A | Option B | Recommended |
|---|---|---|---|
| Context strategy | Read entire codebase | Selective retrieval | Selective (cost, context limit) |
| Planning depth | Plan all steps upfront | Plan one step at a time | Hybrid: high-level plan, detailed per-step |
| Error handling | Give up after N failures | Always retry with feedback | Retry with escalation (ask human after N) |
| Model choice | One model for everything | Specialized models per task | One frontier model (simplicity) |
| Verification | Trust model's assessment | Run tests automatically | Always run tests (trust but verify) |
#The System Design Template
For any AI system design, follow this checklist:
System Design Template
Try it! Pick a product you use daily (Spotify, Netflix, YouTube, Google Search). Set a 10-minute timer and sketch the system design on paper: what data comes in, what model makes predictions, how predictions reach the user, and what could go wrong. Compare your sketch to the case studies below — you will be surprised how much you already understand.
Tests · Calculate daily and monthly costs with model routing. Verify routing saves significant cost vs. single-model. Add caching optimization and compute additional savings.
#Key Takeaways
- Start with requirements, not model architecture. Understand the latency, throughput, accuracy, and cost constraints before choosing models; a recommendation system serving millions needs a fundamentally different architecture than a research tool
- Decompose complex systems into clear components. Every AI system has data ingestion, feature/embedding computation, model inference, serving, and monitoring layers; designing each independently makes the system tractable
- Trade-offs are unavoidable and must be explicit. Latency vs throughput, accuracy vs cost, complexity vs maintainability; document the trade-offs you chose and why, so future engineers understand the design decisions
- Real systems combine multiple ML techniques. A production recommendation system might use collaborative filtering, content-based models, re-ranking, and real-time personalization together, not just one algorithm
#Quick Check
In a recommendation system, why is a two-stage approach (candidate generation + ranking) necessary instead of scoring all items for all users?