Observability for AI Agents: Seeing What Your System Actually Did
Traditional monitoring tells you the service is up. It can't tell you the agent just gave a customer bad advice. Here's how to close that gap.
Why agents break your monitoring
When a traditional service fails, the signals are familiar: error rates spike, latency climbs, a health check goes red. You check logs, metrics, and traces, find the culprit, and fix it.
When an AI agent fails, you often get something stranger: a confidently wrong answer and a 200 OK. Every dashboard is green. Latency is normal. Meanwhile a user just received an incorrect refund policy, a broken SQL query, or an email sent to the wrong person.
The core problem: with agents, correctness is a separate axis from availability. Your system can be perfectly healthy and perfectly wrong.
Agents make this harder than a simple LLM call because their behavior is a chain of decisions:
The model reads context (system prompt, conversation history, retrieved documents)
It decides whether to answer or call a tool
It picks a tool and generates arguments
The tool runs and returns something, maybe an error, maybe garbage
The model interprets that result and decides what to do next
Repeat until it decides it's done (or something stops it)
A failure at step 5 might have been caused by a bad retrieval at step 1. Debugging means reconstructing that entire chain: what the agent saw, what it decided, and why. Without instrumentation, you're left with the final output and a guess.
The anatomy of an agent trace
The good news is that distributed tracing is already the right model. Treat each user request as a trace, and every meaningful step as a span inside it:
Even this simplified view answers questions that a bare log line never could: Why was that slow? (a database query inside a tool.) Why did the agent escalate? (a permissions error on step 6.) How much did this cost?
What to capture on each span
Span type | Attributes worth recording |
|---|---|
LLM call | Model name and version, prompt template version, temperature and key parameters, input/output token counts, finish reason, latency, time-to-first-token (for streaming) |
Tool call | Tool name, arguments, result summary, error type and message, retry count, latency |
Retrieval | Query, number of results, document IDs and scores, latency |
Agent step | Step number, chosen action, remaining budget (steps/tokens/time) |
Whole request | User/session/tenant IDs (pseudonymized), feature or workflow name, total cost, outcome |
Two attributes deserve special emphasis. Prompt template version lets you tie behavior changes to specific prompt releases (see Your Prompts Deserve a CI/CD Pipeline). Session ID lets you stitch together multi-turn conversations, since the real failure often appears three turns after its cause.
Instrumenting with OpenTelemetry
You don't need a proprietary platform to start. OpenTelemetry (OTel) is a strong foundation for three reasons:
AI spans land in the same backend as your existing service traces, so you can see that a slow answer was really a slow database query inside a tool call.
It's vendor-neutral, so you can change backends without re-instrumenting.
It has emerging GenAI semantic conventions: standardized attribute names for models, token usage, and operations. They're still evolving, so pin your instrumentation library versions and expect some renaming over time.
Here's a minimal Python sketch that wraps an LLM call and a tool call:
And the agent loop that ties them together:
A few practical notes:
Wrap your LLM client once, in a single choke point, rather than sprinkling spans across the codebase. It keeps instrumentation consistent and easy to change.
Propagate context across service boundaries, including into tool services and queues, so the trace doesn't break in the middle.
Record the outcome explicitly.
answered,step_limit_hit,escalated, andtool_failureare far more useful in queries than inferring status from logs.Auto-instrumentation libraries exist for popular SDKs and agent frameworks. They're a fast way to start, but review what they capture, especially regarding sensitive content.
The metrics that matter
Beyond the usual request rate, error rate, and latency (RED metrics), agents need a few extra signals.
Cost and usage
Cost per request, computed from token counts and model pricing
Tokens per request (input vs. output), which is often where surprises hide
Cost per successful task, a better measure than cost per call, since failed and retried work still costs money
Cost per tenant or feature, so you can find the expensive outliers
Reliability
Tool-call failure rate, per tool
Retry counts and retry storms
Timeouts and rate-limit (
429) responses from model providersProvider fallback rate, if you fail over between models
Agent behavior
Steps per task, as a distribution (p50, p95, max)
Step-limit hits per hour
Tool-selection distribution, meaning which tools get called and how that shifts after a release
Refusal rate, since sudden changes often signal a prompt or model regression
Quality proxies
You can't measure "correctness" directly at scale, but you can watch signals that correlate with it:
User thumbs up/down ratio
Regeneration rate: how often users ask for another answer
Escalation-to-human rate
Abandonment after an agent response
Follow-up correction rate, such as users re-asking or rephrasing the same question
A sudden shift in any of these after a deploy is a strong hint that something changed, even when latency and errors look normal.
Failure modes worth alerting on
Agents fail in characteristic ways. Build alerts for these specifically:
Failure mode | What it looks like | Signal to alert on |
|---|---|---|
Runaway loop | Agent repeats the same tool call or oscillates between steps | Steps per task above threshold; repeated identical tool calls |
Cost blowout | One request or tenant burns huge token volume | Cost per request p99; tenant-level spend anomalies |
Tool failure cascade | A downstream API degrades and the agent keeps retrying | Tool error rate by tool; retry counts |
Silent regression | A prompt or model change degrades quality without errors | Quality proxies vs. baseline; eval scores on sampled traffic |
Context overflow | Inputs exceed the window, so content is silently truncated | Input token counts near limit; truncation flags |
Hallucinated tool args | Model invents IDs or parameters that don't exist | Validation errors, "not found" rates on tool calls |
Prompt injection | Retrieved or user content hijacks the agent's behavior | Unexpected tool use; policy-violation classifier hits |
The loop metric deserves special attention, because it's the failure that hurts your budget fastest. A stuck agent can make dozens of model calls in a minute. Pair your alert with a hard circuit breaker: a maximum number of steps, tokens, and wall-clock seconds per request, enforced in code. Alerts tell you about the problem, and limits stop it.
Handling sensitive data
Agent traces are uniquely dangerous to store. Prompts routinely contain personal data, credentials, medical or financial details, and proprietary documents, and tool arguments and results can be even more sensitive. Observability that creates a compliance incident isn't an improvement.
Decide these things deliberately, before you ship:
What gets stored raw? For many flows, metadata alone (token counts, latency, tool names, outcome) is enough.
Redaction: scrub known patterns (emails, phone numbers, API keys, card numbers) before data leaves your service, not afterward in the backend. Do it in a span processor or a collector pipeline.
Hashing: log a hash of tool arguments to detect repeats and loops without storing the values themselves.
Access control: restrict who can view full prompt and response content, and audit that access.
Retention: short by default. Keep raw content for days or weeks, aggregated metrics for longer.
Tenant isolation: ensure one customer's data can't surface in another's debug view.
Opt-in capture: make full-content capture a configuration switch you can turn on for a specific session while debugging, rather than an always-on default.
Regional and legal requirements: check where trace data is stored and whether deletion requests (for example, under privacy regulations) can be honored across your telemetry systems.
Involve your security and privacy teams early. It's much easier to design this in than to retrofit it.
Sampling without going blind
Storing every full trace with complete prompts is expensive, both in dollars and in risk. But naive random sampling throws away exactly the traces you most want.
Use tail-based sampling, which decides what to keep after the trace completes:
Keep 100% of traces with errors, tool failures, or step-limit hits
Keep 100% of slow requests (say, above p95 latency)
Keep 100% of traces with negative user feedback or human escalation
Keep 100% of traces from new prompt versions during a canary
Sample a small percentage (1 to 10%) of everything else as a baseline
The OpenTelemetry Collector supports tail-sampling policies for this. You retain the interesting cases at a fraction of the cost, and you always have a representative slice of normal traffic for comparison.
A useful separation: keep lightweight metrics for 100% of traffic (cheap, aggregated, no content) and rich traces for the interesting subset.
Measuring quality in production
Metrics and proxies tell you something changed. To know whether answers are actually good, add evaluation on live traffic:
Sample-based automated grading. Run an LLM judge (calibrated against human labels) over a random sample of production traces, scoring dimensions such as groundedness, correctness of tool use, and policy compliance. Track the scores as time series and alert on drops.
Deterministic checks on every response. Schema validity, presence of required citations, absence of forbidden content. They're cheap enough to run on everything.
Human review queues. Route low-confidence, negatively rated, or randomly sampled traces to reviewers. Give them a fast interface showing the full trace, not just the final answer.
Segment your results. Quality often differs by language, tenant, topic, or input length. An average can hide a serious problem in one segment.
Compare across versions. Tag every trace with prompt and model versions so you can compare a new release against the previous one on live data.
Remember that judges have biases and can drift, so periodically re-check them against human decisions.
Close the loop
The most valuable thing you can do with traces is feed them back into development.
When a user reports a bad answer, you should be able to:
Find the exact trace from a session or request ID
See where the reasoning went wrong: bad retrieval, wrong tool choice, hallucinated argument, misread tool output
Fix the cause: a prompt change, a better tool description, a validation step, or a guardrail
Turn the trace into a regression test in your eval set, with sensitive data scrubbed
Ship the fix through your CI pipeline and watch the same metric in production
Production failures become tomorrow's test cases, and the system gets measurably more reliable with use. This loop is the real payoff of observability: it isn't a dashboard, it's a learning system.
Start simple: a rollout plan
You don't need a specialized platform on day one. A practical order of operations:
Week 1: Basic tracing
Wrap your LLM client and tool executor with spans
Record model, prompt version, token counts, latency, tool name, and outcome
View traces in whatever backend you already run
Week 2: Metrics and guardrails
Add cost per request, tool failure rate, and steps per task
Enforce hard limits on steps, tokens, and time
Alert on step-limit hits and cost outliers
Week 3: Data safety and sampling
Implement redaction and retention rules
Configure tail-based sampling to keep errors, slow, and unhappy-user traces
Week 4: Quality and feedback
Capture thumbs up/down and escalation events on the trace
Start a small sample-based judge, calibrated with human labels
Build the habit of turning bad traces into eval cases
Later, once you know which questions you keep asking, consider dedicated LLM observability tooling for prompt playgrounds, session replay, and dataset management. By then you'll know what you actually need.
Checklist
Every request is a trace; every LLM and tool call is a span
Model, prompt version, and token counts are recorded on LLM spans
Tool name, outcome, and errors are recorded on tool spans
Traces from different services are connected via context propagation
Cost per request and cost per successful task are tracked
Hard limits on steps, tokens, and wall-clock time are enforced in code
Alerts exist for loops, cost spikes, and tool failure cascades
Sensitive data is redacted before export, with defined retention and access rules
Tail-based sampling keeps every error, slow, and negatively rated trace
User feedback is attached to traces
A sampled judge (calibrated against humans) scores production quality
Bad traces routinely become eval cases
Closing thought
Agents don't fail like servers do. They fail like people: plausibly, confidently, and for reasons that only make sense once you see what they saw and what they tried. Observability is how you get to see that. Instrument the chain of decisions, watch cost and behavior alongside latency, protect sensitive data, and feed every failure back into your tests. The teams that do this find out about problems from their dashboards, not their customers.
Previous: The Real Cost of GPU Inference on Kubernetes · Start of series: Your Prompts Deserve a CI/CD Pipeline