Agent Observability: Tracing, Debugging, and Monitoring in Production
After this lesson, you will be able to:
- Add OpenTelemetry tracing to your agents so every LLM call and tool call is captured in a visual tree you can inspect
- Read a LangSmith trace to diagnose exactly where an agent failed, which tool call went wrong, which argument was bad, how many tokens were wasted
- Build a monitoring dashboard that tracks cost and speed for every agent run, so cost surprises never happen again
- Thread a trace ID through multi-agent handoffs so you can follow one task across multiple agents in your logs
- Set up token budget alerts that warn you (or auto-kill the agent) before a runaway loop burns through your API credits
Before You Start
Observability is your X-ray vision into what your agent is actually doing. Without it, your agent is a black box — you send in a question and get back an answer, with no idea what happened in between. This lesson gives you the tools to see everything.
#The Crash That Woke No One Up
Try it! Next time you use an AI chatbot, count how many "steps" it seems to take for a complex query. Now imagine each step had a timestamp, token count, and cost. That is what tracing gives you. If you have access to LangSmith (free tier available), try running a simple LangChain agent and clicking on the trace — the visual tree of every step is incredibly illuminating.
Agents fail silently. An LLM returns JSON with a wrong field name. The tool fails with a validation error. The agent retries with a slightly different argument, and the retry works. Your user sees the correct answer. Your monitoring sees "success." But somewhere in the logs, hidden across 12 tool calls, is a $0.60 retry spiral that should have cost $0.12.
#The Observability Stack for Agents
Production observability requires four concepts working together. Understanding each one independently makes the whole system click.
#Traces: The Execution Tree
trace_id and see exactly what the agent received, decided, and executed.#Spans: Individual Operations
- Name: What operation this is (e.g., "anthropic.messages.create", "tool_call.search_web")
- Start and end timestamps: Used to calculate latency
- Attributes: Metadata specific to this operation (model name, temperature, input token count, output token count)
- Status: Success, error, or timeout
- Events: Notable moments within the span (cache hit, retry triggered, fallback activated)
Spans nest to form the trace tree. An LLM call span might have child spans for the token streaming events, or for a tool call the LLM decided to make during that response.
#Metadata: The Numbers That Matter
Every span should capture structured metadata. For LLM call spans:
{
"model": "claude-opus-4-5",
"input_tokens": 2847,
"output_tokens": 312,
"latency_ms": 1840,
"temperature": 0.7,
"stop_reason": "tool_use",
"run_id": "run_a3b7c912",
"session_id": "sess_user_4821"
}
For tool call spans:
{
"tool_name": "search_web",
"argument_query": "current bitcoin price USD",
"result_tokens": 420,
"tool_latency_ms": 890,
"tool_status": "success",
"retry_count": 0
}
#Events: Notable Moments
Events are lightweight markers within a span. They do not have their own duration — they are just timestamps with a label. Use events to mark moments like:
"retry.triggered"— the tool failed and the agent is retrying"cache.hit"— a cached tool result was used instead of a fresh API call"fallback.activated"— the primary tool failed, switching to backup"context.truncated"— the conversation history was trimmed to fit context window
Events let you ask questions like "how many retries happened across all runs last week?" without having to parse through full span logs.
#OpenTelemetry for Agents
OpenTelemetry (OTel) is the vendor-neutral standard for instrumentation. Once you instrument your agent with OTel spans, you can send the data to any backend: LangSmith, Arize Phoenix, Datadog, Jaeger, Honeycomb, or your own Grafana stack.
#Manual Instrumentation
The most explicit approach — wrap each operation with a span:
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import ConsoleSpanExporter, BatchSpanProcessor
import anthropic
import json
# Set up the tracer
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("my-agent")
client = anthropic.Anthropic()
def run_agent(user_message: str, tools: list) -> str:
"""Run an agent loop with full OTel instrumentation."""
with tracer.start_as_current_span("agent.run") as root_span:
root_span.set_attribute("agent.user_message", user_message)
root_span.set_attribute("agent.tool_count", len(tools))
messages = [{"role": "user", "content": user_message}]
total_input_tokens = 0
total_output_tokens = 0
step = 0
while True:
step += 1
# Span for each LLM call
with tracer.start_as_current_span(f"llm.call.step_{step}") as llm_span:
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=1024,
tools=tools,
messages=messages
)
# Capture token metadata
llm_span.set_attribute("llm.input_tokens", response.usage.input_tokens)
llm_span.set_attribute("llm.output_tokens", response.usage.output_tokens)
llm_span.set_attribute("llm.stop_reason", response.stop_reason)
llm_span.set_attribute("llm.step", step)
total_input_tokens += response.usage.input_tokens
total_output_tokens += response.usage.output_tokens
if response.stop_reason == "end_turn":
root_span.set_attribute("agent.total_input_tokens", total_input_tokens)
root_span.set_attribute("agent.total_output_tokens", total_output_tokens)
root_span.set_attribute("agent.steps", step)
# Estimate cost (Claude claude-opus-4-5 pricing as example)
cost = (total_input_tokens * 0.000015) + (total_output_tokens * 0.000075)
root_span.set_attribute("agent.estimated_cost_usd", round(cost, 6))
return response.content[0].text
# Process tool calls
messages.append({"role": "assistant", "content": response.content})
tool_results = []
for content_block in response.content:
if content_block.type == "tool_use":
with tracer.start_as_current_span(f"tool.{content_block.name}") as tool_span:
tool_span.set_attribute("tool.name", content_block.name)
tool_span.set_attribute("tool.input", json.dumps(content_block.input))
try:
result = execute_tool(content_block.name, content_block.input)
tool_span.set_attribute("tool.result_length", len(str(result)))
tool_span.set_attribute("tool.status", "success")
except Exception as e:
tool_span.set_attribute("tool.status", "error")
tool_span.set_attribute("tool.error", str(e))
tool_span.add_event("tool.error.occurred")
result = {"error": type(e).__name__, "message": str(e)}
tool_results.append({
"type": "tool_result",
"tool_use_id": content_block.id,
"content": json.dumps(result)
})
messages.append({"role": "user", "content": tool_results})#Auto-Instrumentation
openinference library provides auto-instrumentation for Anthropic, OpenAI, LangChain, and more:from openinference.instrumentation.anthropic import AnthropicInstrumentor
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor
# Send traces to Phoenix (self-hosted) or any OTLP-compatible backend
exporter = OTLPSpanExporter(endpoint="http://localhost:6006/v1/traces")
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(exporter))
# This patches anthropic.Anthropic() at import time
AnthropicInstrumentor().instrument(tracer_provider=provider)
# From here on, every client.messages.create() call is automatically traced
# No code changes needed in your agent logicAuto-instrumentation is the recommended approach for existing codebases. You get traces for every API call with zero changes to your agent code.
#LangSmith and LangFuse: Trace UIs
Once traces are flowing, you need a UI to read them. LangSmith (by LangChain) and LangFuse (open source) are the two dominant options.
#What a Trace Looks Like in the UI
Open any trace in LangSmith and you see:
-
Waterfall chart: Each span as a horizontal bar, showing start time and duration. LLM calls are wide (1-3 seconds). Tool calls vary (50ms for cache hits, 2+ seconds for slow APIs). Retries show as duplicate spans at the same nesting level.
-
Token cost breakdown: Each LLM call span shows input tokens, output tokens, and estimated cost. The root span aggregates totals. You can see at a glance which step consumed the most tokens.
-
Input/output viewer: Click any span to see the exact text sent to the model and the exact text returned. This is where you find the "wrong tool argument" bug — you see the model output
{"tool": "search_products", "arguments": {"query": "order_id: 12345"}}when it should have calledlookup_order. -
Latency heatmap: Across many runs, which spans are consistently slow? Which tool calls have high P95 latency?
#Adding Evaluators to Traces
LangSmith supports attaching LLM-as-judge evaluators to traces. After each run, a second LLM call automatically scores the output:
from langsmith import Client
client = Client()
# Create an evaluator that scores tool call efficiency
def efficiency_evaluator(run, example):
tool_call_count = sum(
1 for step in run.child_runs
if step.run_type == "tool"
)
# Flag runs that used more than 10 tool calls as inefficient
score = 1.0 if tool_call_count <= 10 else max(0, 1.0 - (tool_call_count - 10) * 0.1)
return {"key": "tool_efficiency", "score": score}
# Attach to a dataset and run evaluations
client.evaluate(
target=run_agent,
data="my-agent-test-dataset",
evaluators=[efficiency_evaluator],
experiment_prefix="v2-prompt-update"
)#Filtering Traces
In production, you will have thousands of traces. Use tags and metadata to filter:
run_id: Find a specific execution (log the run_id when a user reports a bug)session_id: Find all traces for a specific user sessiontags: Label traces by feature ("flight-booking", "code-generation")latency > 10s: Find slow outlierscost > $0.50: Find expensive runsstatus = error: Find all failed runs
An agent that books flights costs $0.02 per run on average but occasionally costs $0.80. What should you add to your tracing to find the root cause?
retry_count as a span attribute on every tool call span. Track input_tokens per LLM call — a $0.80 run likely has a context window inflation event where a large tool result was not truncated. Add a context_length attribute at each step so you can see the conversation growing. Filter traces by cost > $0.20, then look for spans with retry_count > 0 or LLM call spans with unusually high input_tokens.#Debugging a Real Agent Failure
#The Run That Failed
run_4f7a2b. You open it in LangSmith.#Step 1 — LLM Decides to Search
tool_use block: {"tool": "search_flights", "arguments": {"origin": "SFO", "destination": "JFK", "departure": "2025-03-15", "return": "2025-03-22", "max_price": 400}}. Arguments look correct. Tool call span follows — status: success. Latency: 2.1 seconds. Result: 8 flight options returned, 3,200 tokens of JSON.#Step 2 — Wrong Argument Name on Book
{"flight_id": "UA-4821", "passenger_name": "...", "card_num": "..."}. The actual tool schema requires "card_number", not "card_num". The tool call span shows status: error, error: "ValidationError: 'card_num' is not a valid field".#Step 3 — Retry Spiral Begins
card_num. The tool description says card_number but the model trained on a slightly different schema. Three retries, all with the same wrong field name. Each retry: 1 LLM call + 1 failed tool call = ~2,000 tokens. Three retries = 6,000 extra tokens.#Step 4 — Context Overflow
By the fourth LLM call, the context window has grown: the original search results (3,200 tokens) + the booking attempts + the error messages. Input token count on this LLM call: 12,400. The model is now reasoning about all three previous failures simultaneously. It starts hedging and eventually generates: "I was unable to complete your booking."
#The Fix
card_number explicitly with an example. (2) Add Pydantic validation before the tool executes — catch the card_num vs card_number mismatch and return a structured error with the correct field name: {"error": "WrongFieldName", "expected": "card_number", "received": "card_num"}. This gives the model a clear correction signal instead of a generic ValidationError.#Distributed Tracing for Multi-Agent Systems
Single-agent tracing is straightforward. Multi-agent tracing requires explicit context propagation — the trace_id must travel with every message.
#Propagating Trace Context
from opentelemetry import trace, context
from opentelemetry.propagate import inject, extract
import asyncio
tracer = trace.get_tracer("multi-agent")
async def orchestrator(task: str):
with tracer.start_as_current_span("orchestrator.run") as root_span:
root_span.set_attribute("task", task)
# Inject trace context into carrier dict to pass to sub-agents
carrier = {}
inject(carrier) # carrier now contains trace_id and span_id
# Spawn sub-agents in parallel, passing the trace context
results = await asyncio.gather(
sub_agent("research", task, carrier),
sub_agent("planning", task, carrier),
sub_agent("execution", task, carrier),
)
root_span.set_attribute("sub_agent_count", 3)
return results
async def sub_agent(name: str, task: str, parent_carrier: dict):
# Extract parent context from carrier — links this span as a child
parent_context = extract(parent_carrier)
with tracer.start_as_current_span(
f"sub_agent.{name}",
context=parent_context
) as span:
span.set_attribute("sub_agent.name", name)
span.set_attribute("sub_agent.task", task)
# This sub-agent's LLM calls and tool calls will nest here
result = await run_sub_agent_loop(name, task)
span.set_attribute("sub_agent.result_length", len(result))
return resultorchestrator.run → 3 children (sub_agent.research, sub_agent.planning, sub_agent.execution), each with their own LLM call and tool call children. You can see all 4 agents' work in one unified waterfall.#The trace_id: Your Single Key
trace_id is the most important string in production agent debugging. It is the single key that lets you reconstruct any execution:- Log it everywhere: in your application logs, in your error tracker (Sentry), in your database with the task record
- Include it in user-facing error messages: "Error reference: run_4f7a2b" — users can report it, you can find the exact trace
- Use it for cost attribution: aggregate total cost by trace_id to know exactly what each user task cost
A multi-agent system: orchestrator spawns 3 sub-agents in parallel. Sub-agent #2 fails. How do you correlate its logs with the orchestrator's trace?
The answer is context propagation. Passing the parent carrier (containing trace_id and span_id) to each sub-agent at spawn time means all sub-agents' spans appear as children under the orchestrator's span in the same trace. You open one trace in LangSmith and see the full picture: orchestrator → sub-agent-1 (success) → sub-agent-2 (error, 3 retries) → sub-agent-3 (success, waiting for sub-agent-2). The failure is immediately visible.
#Cost Monitoring
Token costs are not uniform. The same agent task can cost $0.02 on a fast day and $0.80 on a bad one. Cost monitoring tells you which tasks, which tools, and which users are expensive.
#Per-Run Cost Estimation
Calculate cost at the span level, aggregate at the trace level:
# Cost rates per 1M tokens (example rates, verify current pricing)
COST_PER_MILLION = {
"claude-opus-4-5": {"input": 15.00, "output": 75.00},
"claude-sonnet-4-5": {"input": 3.00, "output": 15.00},
"claude-haiku-3-5": {"input": 0.80, "output": 4.00},
}
def calculate_span_cost(model: str, input_tokens: int, output_tokens: int) -> float:
"""Calculate the cost for a single LLM call span."""
if model not in COST_PER_MILLION:
return 0.0
rates = COST_PER_MILLION[model]
input_cost = (input_tokens / 1_000_000) * rates["input"]
output_cost = (output_tokens / 1_000_000) * rates["output"]
return round(input_cost + output_cost, 6)
def add_cost_to_span(span, model: str, usage):
"""Add cost attributes to an LLM call span."""
cost = calculate_span_cost(model, usage.input_tokens, usage.output_tokens)
span.set_attribute("cost.input_tokens", usage.input_tokens)
span.set_attribute("cost.output_tokens", usage.output_tokens)
span.set_attribute("cost.usd", cost)
span.set_attribute("cost.model", model)
return cost#Token Budget Alerts
Kill expensive runs before they spiral. Check token count at each step and abort if the budget is exceeded:
MAX_TOKENS_PER_RUN = 50_000 # $0.75 at claude-opus-4-5 rates
MAX_STEPS = 20
class TokenBudgetExceeded(Exception):
pass
class StepLimitExceeded(Exception):
pass
def run_agent_with_budget(user_message: str, tools: list) -> str:
total_tokens = 0
steps = 0
with tracer.start_as_current_span("agent.run") as root_span:
while True:
steps += 1
if steps > MAX_STEPS:
root_span.set_attribute("agent.abort_reason", "step_limit")
root_span.add_event("budget.step_limit_exceeded")
raise StepLimitExceeded(f"Agent exceeded {MAX_STEPS} steps")
response = client.messages.create(...)
total_tokens += response.usage.input_tokens + response.usage.output_tokens
if total_tokens > MAX_TOKENS_PER_RUN:
root_span.set_attribute("agent.abort_reason", "token_budget")
root_span.add_event("budget.token_limit_exceeded")
raise TokenBudgetExceeded(f"Token budget exceeded: {total_tokens} > {MAX_TOKENS_PER_RUN}")
if response.stop_reason == "end_turn":
root_span.set_attribute("agent.final_tokens", total_tokens)
return response.content[0].text#Cost Attribution
Track cost by the dimensions that matter for your business:
| Dimension | Why It Matters |
|---|---|
| Per tool call | Which tool is the most expensive to call? (Often: the LLM calls that process large tool results) |
| Per agent type | Is the research agent 5× more expensive than the execution agent? |
| Per user session | Which users generate the most cost? (Rate limiting candidates) |
| Per task category | Is "flight booking" 10× more expensive than "FAQ lookup"? (Pricing decisions) |
#Latency SLOs for Agents
Agents are slow. An agent that takes 45 seconds for a task users expect in 5 is a product problem, not just a performance problem.
#Measuring Agent Latency
import time
from dataclasses import dataclass
from typing import List
@dataclass
class StepLatency:
step: int
operation: str
latency_ms: float
def measure_agent_latency(agent_func, user_message: str, n_runs: int = 20):
"""Measure P50/P95/P99 latency across multiple runs."""
run_times = []
step_times: List[StepLatency] = []
for _ in range(n_runs):
start = time.time()
agent_func(user_message)
end = time.time()
run_times.append((end - start) * 1000) # milliseconds
run_times.sort()
p50 = run_times[int(n_runs * 0.50)]
p95 = run_times[int(n_runs * 0.95)]
p99 = run_times[int(n_runs * 0.99)]
print(f"Latency (ms): P50={p50:.0f} P95={p95:.0f} P99={p99:.0f}")
return {"p50": p50, "p95": p95, "p99": p99}#Slow Tool Identification
When your P95 latency is unacceptable, find the bottleneck. Sort tool call spans by average latency:
from collections import defaultdict
import statistics
def analyze_tool_latency(traces: list) -> dict:
"""Find the slowest tools across all traces."""
tool_latencies = defaultdict(list)
for trace in traces:
for span in trace.spans:
if span.name.startswith("tool."):
tool_name = span.attributes.get("tool.name")
latency_ms = (span.end_time - span.start_time) / 1_000_000 # ns to ms
tool_latencies[tool_name].append(latency_ms)
results = {}
for tool, latencies in tool_latencies.items():
latencies.sort()
n = len(latencies)
results[tool] = {
"p50": latencies[int(n * 0.5)],
"p95": latencies[int(n * 0.95)],
"call_count": n,
"avg": statistics.mean(latencies),
}
# Sort by P95 descending to find worst offenders
return dict(sorted(results.items(), key=lambda x: x[1]["p95"], reverse=True))#Production Dashboard Pattern
What to track per agent deployment. These 7 metrics constitute a minimal viable observability dashboard:
| Metric | What to Measure | Alert Threshold |
|---|---|---|
| Success rate | % of runs that complete without error | Alert if drops below 95% |
| P95 latency | 95th percentile end-to-end run time | Alert if exceeds SLO (e.g., 30s) |
| Cost per run | Average tokens × price across last 100 runs | Alert if spikes 3× baseline |
| Token budget hit rate | % of runs that hit the token budget limit | Alert if above 5% (tuning signal) |
| Retry rate | Average retry events per run | Alert if above 0.5 (tool schema quality signal) |
| Context overflow rate | % of runs where context exceeded 80% of max | Alert if above 10% (chunking needed) |
| Tool error rate | % of tool calls that return errors | Alert if any tool exceeds 5% error rate |
#Putting It All Together
Tests · The root span should have estimated_cost_usd, total_input_tokens, and total_output_tokens attributes. Each tool call should have a tool.name attribute and tool.status='success'. The trace should have at least 5 spans total.
#Key Takeaways
- Traces are the unit of agent debugging. A trace captures the complete execution tree of one agent run; every LLM call, every tool call, every retry is a span in that tree; the trace_id is the single key to reconstruct any failure
- Cost surprises require per-span token tracking. Aggregate token counts at run completion misses retry spirals; tracking input_tokens and output_tokens per LLM call span reveals exactly which step caused a cost spike
- Distributed tracing requires explicit context propagation. When spawning sub-agents, inject the parent trace context into the message; the sub-agent extracts it and creates child spans; without this, multi-agent systems produce disconnected log streams you cannot correlate
- Seven metrics constitute a minimal production dashboard. Success rate, P95 latency, cost per run, token budget hit rate, retry rate, context overflow rate, and tool error rate; alert on any that cross a baseline threshold
#Quick Check
What is the primary purpose of a span's 'events' in OpenTelemetry?