LLMOps: Operating Large Language Models
After this lesson, you will be able to:
- Explain how running LLMs in production is different from traditional ML, and why you need new tools for it
- Set up prompt versioning and testing so that changing a prompt is as safe and traceable as changing code
- Build an LLM gateway that routes requests, limits costs, tracks spending, and switches providers when one goes down
Before You Start
#MLOps vs LLMOps: Same Principles, Different World
If you have used ChatGPT or any LLM API, you have already experienced LLMOps challenges without knowing it: inconsistent responses, surprise costs, prompt tweaks that break things. This lesson gives you the tools and patterns to manage all of that professionally.
Your LLM endpoint costs $50K/month. What is the first optimization to try?
The answer is D. Before optimizing anything, you need to understand what you are spending on. Traffic analysis often reveals that 20-40% of LLM calls are unnecessary — queries that could be handled by keyword search, rule engines, or cached responses. The cheapest LLM call is the one you never make.
#The LLMOps Stack
Traditional MLOps has experiment tracking, model registries, and CI/CD. LLMOps needs all of that plus five additional layers:
#1. Prompt Management
Prompts are the "code" of LLM applications. They need the same rigor as software code: version control, review, testing, and rollback.
| Tool | Type | Key Feature |
|---|---|---|
| Humanloop | SaaS platform | Visual prompt editor + A/B testing |
| PromptLayer | SaaS platform | Prompt version history + analytics |
| Braintrust | SaaS platform | Eval-first prompt development |
| Langfuse | Open source | Prompt management + observability |
| Git + JSON files | DIY | Free, full control, integrates with existing CI |
A production-grade prompt versioning workflow
This is exactly how software CI/CD works — applied to prompts instead of code.
#2. Guardrails Infrastructure
Guardrails are the safety nets that prevent LLMs from going off the rails. They sit between your application and the LLM, intercepting both inputs and outputs.
Input guardrails (before the LLM sees the query)
- Prompt injection detection (is the user trying to override the system prompt?)
- PII detection and redaction (remove credit card numbers, SSNs before processing)
- Topic filtering (is this query within the allowed scope?)
- Token budget enforcement (reject queries that would cost too much)
Output guardrails (before the user sees the response)
- Hallucination detection (does the response contain claims not supported by context?)
- Content safety filtering (toxic, harmful, or inappropriate content)
- Format validation (is the JSON output actually valid JSON?)
- Factual grounding checks (are cited facts accurate?)
Key guardrails tools
| Tool | Approach | Best For |
|---|---|---|
| NVIDIA NeMo Guardrails | Programmable rails in Colang | Complex, custom guardrail logic |
| Guardrails AI | Python validators with Pydantic | Structured output validation |
| LLM Guard | Open source input/output scanner | Quick safety layer |
| Lakera Guard | API-based prompt injection detection | Prompt security specifically |
Try it! Take any LLM API (or use a free one) and try the same prompt 5 times. Notice how the outputs differ each time — same input, different output. Now imagine versioning and testing those prompts automatically. That is what LLMOps tools do.
# Example: Guardrails AI for structured output validation
from guardrails import Guard
from guardrails.hub import ValidRange, ToxicLanguage
guard = Guard().use_many(
ValidRange(min=0, max=100, on_fail="fix"),
ToxicLanguage(on_fail="filter"),
)
# The guard wraps your LLM call and validates the output
raw_response, validated_response, *rest = guard(
llm_api=openai.chat.completions.create,
model="gpt-4o",
messages=[{"role": "user", "content": "Rate this product 1-100"}],
)
# If the LLM returns 150, the guard fixes it to 100
# If the LLM outputs toxic language, the guard filters it#3. LLM Gateways and Routers
An LLM gateway sits between your application and LLM providers, providing a unified interface regardless of which model or provider you use.
What a gateway does
- Unified API: Call OpenAI, Anthropic, Google, or self-hosted models through one interface
- Load balancing: Distribute requests across multiple API keys or providers
- Failover: If OpenAI is down, automatically route to Anthropic
- Rate limiting: Prevent any single user or feature from consuming the entire token budget
- Cost tracking: Monitor spend per team, per feature, per user
- Caching: Semantic caching at the gateway level
Key gateway tools (2024-2026 landscape)
| Tool | Type | Differentiator |
|---|---|---|
| LiteLLM | Open source proxy | 100+ model providers, drop-in OpenAI replacement, built-in budgets |
| Portkey | SaaS gateway | Production-grade fallback chains, caching, observability |
| Helicone | Open source proxy | Drop-in OpenAI proxy with cost analytics + retries |
| Cloudflare AI Gateway | Managed (edge) | Edge caching + rate limiting + analytics; minimal latency overhead |
| Martian | SaaS router | AI-powered model selection per query |
| OpenRouter | Marketplace | Arbitrage across providers, useful for latency-tolerant traffic |
| Kong AI Gateway | Open source | Enterprise API gateway with AI extensions |
| AWS Bedrock / Azure AI Foundry | Managed (cloud-native) | VPC + IAM + provisioned throughput |
# LiteLLM: unified interface to any LLM provider
from litellm import completion
# Same function, different providers — swap models without code changes
response_openai = completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Explain LLMOps"}],
)
response_anthropic = completion(
model="claude-3-5-sonnet-20241022",
messages=[{"role": "user", "content": "Explain LLMOps"}],
)
response_local = completion(
model="ollama/llama3",
messages=[{"role": "user", "content": "Explain LLMOps"}],
)The math is worth running explicitly — most teams underestimate the savings.
#4. Evaluation Pipelines
LLM evaluation is fundamentally harder than classical ML evaluation. You cannot just compute accuracy on a test set because there is no single "correct" answer for most LLM tasks.
Evaluation dimensions
| Dimension | What It Measures | How to Measure |
|---|---|---|
| Correctness | Is the answer factually right? | Human review, LLM-as-judge |
| Relevance | Does it address the question? | Semantic similarity, LLM-as-judge |
| Faithfulness | Is it grounded in provided context? | NLI models, citation checking |
| Helpfulness | Is it useful to the end user? | User feedback, A/B tests |
| Safety | Is it free of harmful content? | Content classifiers, red teaming |
| Cost | How much did it cost per response? | Token counting, gateway metrics |
| Latency | How fast is the response? | Time-to-first-token, total time |
Key evaluation tools
| Tool | Type | Best For |
|---|---|---|
| promptfoo | Open source CLI | Prompt regression testing in CI |
| LangSmith | SaaS platform | Tracing + evaluation for LangChain apps |
| Braintrust | SaaS platform | Eval datasets + scoring + experiments |
| Ragas | Open source | RAG-specific evaluation metrics |
| DeepEval | Open source | Unit testing framework for LLMs |
Tests · Run both prompt versions against the test suite. Verify the new prompt handles off-topic queries better. Check the deployment decision logic.
#5. Cost Management for LLM Workloads
LLM costs behave nothing like traditional infrastructure costs. A web server at 10x traffic costs ~10x. An LLM application at 10x traffic can cost 10x more per user because each request involves heavy computation.
The cost management stack
- Track cost per request, per user, per feature, per team
- Use gateway-level metrics (LiteLLM, Portkey) for real-time cost dashboards
- Alert on anomalies: cost spikes, token count drift, unusual traffic patterns
- Model routing: cheap models for easy queries, expensive models for hard ones
- Prompt compression: reduce input tokens by 40-60% without quality loss
- Semantic caching: 15-25% cache hit rate for repeated queries
- Batch APIs: 50% cheaper for non-real-time workloads (OpenAI Batch API)
- Set per-team and per-feature token budgets
- Implement hard stops: when a feature hits its monthly budget, degrade gracefully
- Rate limit aggressive users or features to prevent runaway costs
Real-World Cost Reduction Timeline
| Month | Cost | Optimization Applied |
|---|---|---|
| Month 1 | $50,000 | No optimization (baseline) |
| Month 2 | $35,000 | Switch 60% of traffic to GPT-5-mini / Haiku 4.5 |
| Month 3 | $28,000 | Add semantic caching (20% hit rate) |
| Month 4 | $22,000 | Prompt compression (30% fewer input tokens) |
| Month 5 | $18,000 | Batch API for async workloads |
| Total | 64% reduction | Same quality, same features |
#Putting It All Together: The LLMOps Architecture
A production LLMOps stack looks like this:
#2025-2026 LLMOps Advances
The LLMOps stack is moving as fast as the models. Notable additions since 2024:
- Serving runtimes: vLLM 0.6+, SGLang 0.4+, and TensorRT-LLM 0.18+ ship continuous batching, prefix caching, speculative decoding, FP8 / FP4 inference (B200), and disaggregated prefill/decode by default — what required custom infra in 2024 is one config flag in 2026.
- Memory-first agent runtimes: Letta (formerly MemGPT, 2025) productized self-managed agent memory; CrewAI Studio and LangGraph Studio added visual graph IDEs for multi-agent workflows; Autogen 2.0 standardized the multi-agent contract.
- Optimization frameworks: DSPy 2.5+ (2025) made declarative prompt + program optimization mainstream — define a signature, let DSPy compile it against your eval set.
- Apple Silicon serving: MLX (Apple 2024+) lets you run Qwen 3, Llama 4 Scout, and DeepSeek distills on M-series Macs for prototyping and on-device inference.
- Serverless GPU + edge: Modal, Replicate, Cloudflare Workers AI, and Together / Fireworks / Groq cover the spectrum from pay-per-second GPUs to global edge inference; pick by latency budget and request shape.
- Function calling 2.0: parallel and batched tool calls are now standard across OpenAI, Anthropic, Google, and most open-weights serving stacks — your gateway must handle tool-call concurrency, not just sequential calls.
- Native multimodal output: Claude voice, GPT-4o/GPT-5 audio-out, and Gemini native image-out generate non-text modalities directly. Logging now needs to capture audio and image bytes, not just text.
#Common LLMOps Mistakes
-
Treating prompts like configuration, not code — Prompts change behavior as much as code changes. They need version control, code review, testing, and rollback capabilities.
-
No evaluation before deployment — "The new prompt looks good to me" is not a deployment strategy. Run it against 50+ test cases and compare to the baseline.
-
Single-provider dependency — If 100% of your traffic goes to one LLM provider and they have a 4-hour outage, your product is down for 4 hours. Use a gateway with automatic failover.
-
Ignoring cost until it is too late — Set up cost tracking on day one, not after your first $50K bill. Gateway-level cost attribution is cheap to implement and invaluable for budgeting.
-
Over-engineering guardrails — Start with basic input/output validation. Add sophisticated guardrails (NeMo, custom classifiers) only after you have real data on what goes wrong.
#Quick Check
What is the primary difference between LLMOps and traditional MLOps?
#Key Takeaways
- Prompts are code, not config. Version-control every prompt, run regression tests before deploying changes, and maintain rollback capability; a single bad prompt update can silently degrade your entire application
- Guardrails protect both directions. Input guardrails catch prompt injection and PII before the LLM processes them; output guardrails catch hallucination and harmful content before users see them; skip either side and you have a half-protected system
- LLM gateways eliminate provider lock-in. A single unified API (LiteLLM, Portkey) gives you automatic failover, per-team cost tracking, rate limiting, and semantic caching without changing application code
- Evaluation never stops. Run promptfoo or LangSmith in CI to catch regressions before production; track correctness, relevance, safety, and cost as separate dimensions, not a single score
- Prompt caching is the highest-ROI quick win. For RAG systems where the system prompt is identical across queries, Anthropic prompt caching can reduce input token costs by 90% and time-to-first-token by 50%