Cost Architecture
After this lesson, you will be able to:
- Compare GPU options (A100, H100, L4, T4) and pick the right hardware for training versus serving
- Design a routing system that sends easy questions to cheap models and hard questions to expensive ones
- Apply token-saving strategies that cut LLM API costs by 40-70% without making the output worse
Before You Start
#The $100K/Month Surprise
This lesson might save you (or your company) more money than any other topic in this track. Cost surprises kill more AI startups than bad models. The good news: most cost problems are solvable once you know the patterns.
#The Real Numbers: What ML Actually Costs
Tests · Verify annual cost calculation is correct. Confirm serving costs exceed training costs. Calculate the retraining scenario.
#GPU Selection: The Hardware Menu
#Training GPUs
| GPU | VRAM | FP16 TFLOPS | $/hr (cloud) | Best For |
|---|---|---|---|---|
| H100 SXM | 80 GB | 990 | $3.00-4.00 | Large model training (70B+) |
| A100 SXM | 80 GB | 312 | $2.00-2.50 | Most training workloads |
| A100 PCIe | 40/80 GB | 312 | $1.50-2.50 | Budget training |
| L4 | 24 GB | 121 | $0.40-0.60 | Fine-tuning, small model training |
| H200 | 141 GB | 990 | $5.00-6.00 | Very large models (140B+) |
Prices are approximate as of early 2026 and vary by cloud provider and region.
#Inference GPUs
| GPU | VRAM | INT8 TOPS | $/hr (cloud) | Best For |
|---|---|---|---|---|
| H100 | 80 GB | 1,979 | $3.00-4.00 | High-throughput LLM serving |
| A100 | 80 GB | 624 | $2.00-2.50 | Standard LLM serving |
| L4 | 24 GB | 485 | $0.40-0.60 | Cost-efficient inference, small models |
| T4 | 16 GB | 130 | $0.35-0.50 | Budget inference, INT8 models |
| L40S | 48 GB | 733 | $1.00-1.50 | Mid-range, good for 13B-34B models |
Prices are approximate as of early 2026 and vary by cloud provider and region.
You need to serve a Llama 70B model. An A100-80GB costs $2/hr and an L4-24GB costs $0.60/hr. The 70B model in INT4 needs ~37GB. Which is more cost-effective?
This is a trick question — two L4s only have 48 GB total, and tensor parallelism overhead means you likely need the A100. But the real answer is D: cost-effectiveness depends on throughput. If the A100 serves 30 requests/second and two L4s serve 15 req/s, the per-request cost is identical. You need to benchmark your specific model and traffic pattern.
#The Cost Architecture Decision Process
Building a cost-efficient AI system is not about being cheap — it is about being systematic. Here is the step-by-step process for designing an architecture that balances performance and cost:
#Step 1: Estimate Request Volume
Start with your traffic projections. How many requests per day? What is the peak-to-average ratio? A chatbot serving 10K requests/day has fundamentally different economics than one serving 1M. Volume determines whether API providers or self-hosting is more economical.
#Step 2: Calculate API Costs
Price out the API option. At 50K requests/day with 2,000 input tokens and 500 output tokens each, GPT-4-turbo costs ~$30K/month. GPT-5 mid-tier costs ~$7-8K/month. Claude Haiku 4.5 / GPT-5-mini costs ~$1-3K/month. Model choice is the single biggest cost lever.
#Step 3: Calculate Self-Hosted GPU Costs
Price out the self-hosting option. Two A100-80GB GPUs at $2/hr each cost $2,880/month fixed. Add engineering time for setup, monitoring, and maintenance. Self-hosting has a high fixed cost but near-zero marginal cost per request.
#Step 4: Find the Break-Even Point
The break-even point is where API costs equal self-hosted costs. Below this volume, APIs win (simpler, no ops burden). Above it, self-hosting wins (lower marginal cost). For most teams, the break-even is around $15-20K/month in API spend, factoring in engineering time.
#Step 5: Choose Architecture
Select your architecture based on the analysis. Options include: all-API (simple, scales automatically), all-self-hosted (cheapest at scale), or hybrid (self-host a small model for easy queries, API for hard ones). Multi-model routing is almost always the right answer at scale.
#Step 6: Implement Caching and Batching
Layer on cost optimizations. Semantic caching intercepts 15-25% of repeated queries. Prompt compression reduces input tokens by 40-60%. Batch APIs (50% cheaper) handle non-real-time tasks. Each optimization compounds multiplicatively with the others.
#Step 7: Monitor and Optimize
Track cost per request, cost per user, and cost per feature daily. Alert on anomalies (prompt drift making prompts longer, traffic spikes, retry storms). Review cost metrics weekly in sprint reviews. The best cost architecture evolves continuously as traffic patterns and model pricing change.
#Token Optimization: Spending Less on Every Request
For LLM API users (OpenAI, Anthropic, Google), tokens are the primary cost driver:
#1. Prompt Engineering for Efficiency
Before optimization
You are a helpful customer service agent for Acme Corp. You were founded
in 1985 and are known for your excellent products and customer service.
When responding to customers, always be polite, professional, and helpful.
Make sure to address their concern directly and provide actionable steps.
If you don't know the answer, say so honestly rather than making
something up. Always end with asking if there's anything else you can help with.
Customer: What is your return policy?
Token count: ~100 input tokens for system prompt
After optimization
[Acme Corp CS agent. Be concise, helpful, honest. If unsure, say so.]
Customer: What is your return policy?
Token count: ~20 input tokens for system prompt (80% reduction)
#2. Caching Strategies
Prompt caching (provider-level)
- Anthropic and OpenAI offer prompt caching: repeated system prompts are cached server-side
- First request pays full price; subsequent requests with the same prefix get a 90% discount on cached tokens
- Effective for applications with static system prompts
Semantic caching (application-level)
- Hash incoming queries and cache responses
- Use embedding similarity to match similar (not identical) queries to cached responses
- Cache hit rate of 10-30% for customer support bots, higher for FAQ-style products
- Invalidation strategy: TTL-based or event-triggered
#3. Context Window Management
Tests · Calculate the baseline cost and verify each optimization strategy reduces it. Find the break-even point for self-hosting.
#Multi-Model Routing: The Right Model for the Job
Routing strategies
#1. Rule-Based Routing
Rule-Based Routing Logic
Simple but effective for well-understood traffic patterns. Works for 60-70% of use cases.
#2. Classifier-Based Routing
Train a small, fast classifier (BERT-tiny, 10ms latency) to predict query difficulty:
- Class 0: Easy (route to mini model) — $0.001/request
- Class 1: Medium (route to standard model) — $0.01/request
- Class 2: Hard (route to premium model) — $0.05/request
The classifier costs $0.0001/request. Even if it misroutes 10% of queries, the overall cost savings are massive.
#3. Cascade / Fallback Routing
- Try the small model first
- If confidence is low or output quality is poor (detected by a cheap quality check), escalate to the large model
- Only 20-30% of queries typically need escalation
#LLM Gateways: The Cost Control Plane
| Gateway | Open source? | Strongest feature | Notes |
|---|---|---|---|
| LiteLLM | Yes | Unified OpenAI-compatible API across 100+ providers | The de-facto standard for self-hosted gateways; built-in budgets and routing rules. |
| Portkey | Hybrid | Smart routing, fallback chains, prompt versioning | Strong fallback graphs (try GPT-5, on fail try Claude Sonnet 5, on fail try DeepSeek-V3.1). |
| Helicone | Yes | Drop-in OpenAI proxy with cost analytics | Minimal code change; pairs well with Langfuse for tracing. |
| Cloudflare AI Gateway | Managed | Edge caching + rate limiting + observability | Lowest latency overhead; pay only for traffic. |
| OpenRouter | Managed | Marketplace pricing across providers | Useful for arbitrage when latency-tolerant. |
| AWS Bedrock / Azure AI Foundry | Managed | Native cloud auth + VPC + provisioned throughput | Pick if you are already deep in that cloud. |
What you should enforce at the gateway:
- Per-customer / per-feature budgets — hard cutoff plus soft alerts at 80%.
- Prefix caching — hash system prompts and reuse server-side KV; many gateways ship this.
- Semantic cache — embed and match the user message; serve cached completions for high-similarity hits.
- Fallback chains. Primary fails or rate-limits, secondary handles, tertiary handles, and graceful degradation if all fail.
- Provider rotation. When a provider has elevated p99, shift weight automatically.
- Audit logging. Every prompt, completion, model, latency, token count, and cost tagged by user/feature.
Your team is deciding whether to call OpenAI directly from each microservice or to route everything through a self-hosted LiteLLM gateway. Annual LLM spend is projected at $1.2M. What is the single strongest argument for the gateway?
#FinOps Math: Per-Million-Tokens, GPU Hours, and Spot Economics
Pricing models for AI workloads are a mess of dimensions. Work them out in the same units before deciding.
#Spot Instances and Preemptible Compute
Training on spot instances
- AWS Spot, GCP Preemptible, Azure Low Priority: 60-90% cheaper than on-demand
- Risk: instances can be terminated with 30-120 seconds notice
- Mitigation: checkpoint every N minutes, auto-resume from latest checkpoint
- Best practice: Checkpoint every 15-30 minutes. Use a persistent storage volume (EFS, GCS) for checkpoints.
Inference on spot instances
- More risky than training (user-facing latency)
- Use spot for batch inference (non-real-time)
- For real-time: use spot for overflow capacity with on-demand for baseline
#Key Takeaways
- GPU selection depends on workload type. H100s for large-scale training, A100s for medium training and inference, L4/T4 for cost-effective inference; matching hardware to workload avoids 2-5x overspending
- Multi-model routing cuts costs dramatically. Route simple queries to cheap small models and hard queries to expensive large models; 70-80% of queries are often simple enough for the small model
- Token optimization reduces LLM API costs by 40-70%. Prompt compression, caching frequent responses, batching requests, and truncating unnecessary context all reduce token consumption without degrading quality
- Cost must be a first-class architectural concern. Designing for cost-efficiency from the start (choosing the right model size, implementing caching, using spot instances) is far easier than optimizing after launch
#Quick Check
Why is the NVIDIA L4 often the most cost-effective GPU for inference?