Observability for AI Agents: Seeing What Your System Actually Did

Traditional monitoring tells you the service is up. It can't tell you the agent just gave a customer bad advice. Here's how to close that gap.

No headings found on page

Tech

11

Min Read

Share:

Why agents break your monitoring

When a traditional service fails, the signals are familiar: error rates spike, latency climbs, a health check goes red. You check logs, metrics, and traces, find the culprit, and fix it.

When an AI agent fails, you often get something stranger: a confidently wrong answer and a 200 OK. Every dashboard is green. Latency is normal. Meanwhile a user just received an incorrect refund policy, a broken SQL query, or an email sent to the wrong person.

The core problem: with agents, correctness is a separate axis from availability. Your system can be perfectly healthy and perfectly wrong.

Agents make this harder than a simple LLM call because their behavior is a chain of decisions:

  1. The model reads context (system prompt, conversation history, retrieved documents)

  2. It decides whether to answer or call a tool

  3. It picks a tool and generates arguments

  4. The tool runs and returns something, maybe an error, maybe garbage

  5. The model interprets that result and decides what to do next

  6. Repeat until it decides it's done (or something stops it)

A failure at step 5 might have been caused by a bad retrieval at step 1. Debugging means reconstructing that entire chain: what the agent saw, what it decided, and why. Without instrumentation, you're left with the final output and a guess.

The anatomy of an agent trace

The good news is that distributed tracing is already the right model. Treat each user request as a trace, and every meaningful step as a span inside it:

Trace: handle_support_request  (4.8s, $0.021)
├── span: retrieve_context          (0.31s)
│     └── span: vector_search       (0.28s)
├── span: llm_call #1               (1.2s)  → decides to call lookup_order
├── span: tool: lookup_order        (0.9s)
│     └── span: db.query            (0.85s)  ← the actual slow part
├── span: llm_call #2               (1.1s)  → decides to call issue_refund
├── span: tool: issue_refund        (0.2s)  ✗ error: permission_denied
├── span: llm_call #3               (1.0s)  → apologizes, escalates
└── span: escalate_to_human         (0.1s)
Trace: handle_support_request  (4.8s, $0.021)
├── span: retrieve_context          (0.31s)
│     └── span: vector_search       (0.28s)
├── span: llm_call #1               (1.2s)  → decides to call lookup_order
├── span: tool: lookup_order        (0.9s)
│     └── span: db.query            (0.85s)  ← the actual slow part
├── span: llm_call #2               (1.1s)  → decides to call issue_refund
├── span: tool: issue_refund        (0.2s)  ✗ error: permission_denied
├── span: llm_call #3               (1.0s)  → apologizes, escalates
└── span: escalate_to_human         (0.1s)
Trace: handle_support_request  (4.8s, $0.021)
├── span: retrieve_context          (0.31s)
│     └── span: vector_search       (0.28s)
├── span: llm_call #1               (1.2s)  → decides to call lookup_order
├── span: tool: lookup_order        (0.9s)
│     └── span: db.query            (0.85s)  ← the actual slow part
├── span: llm_call #2               (1.1s)  → decides to call issue_refund
├── span: tool: issue_refund        (0.2s)  ✗ error: permission_denied
├── span: llm_call #3               (1.0s)  → apologizes, escalates
└── span: escalate_to_human         (0.1s)

Even this simplified view answers questions that a bare log line never could: Why was that slow? (a database query inside a tool.) Why did the agent escalate? (a permissions error on step 6.) How much did this cost?

What to capture on each span

Span type

Attributes worth recording

LLM call

Model name and version, prompt template version, temperature and key parameters, input/output token counts, finish reason, latency, time-to-first-token (for streaming)

Tool call

Tool name, arguments, result summary, error type and message, retry count, latency

Retrieval

Query, number of results, document IDs and scores, latency

Agent step

Step number, chosen action, remaining budget (steps/tokens/time)

Whole request

User/session/tenant IDs (pseudonymized), feature or workflow name, total cost, outcome

Two attributes deserve special emphasis. Prompt template version lets you tie behavior changes to specific prompt releases (see Your Prompts Deserve a CI/CD Pipeline). Session ID lets you stitch together multi-turn conversations, since the real failure often appears three turns after its cause.

Instrumenting with OpenTelemetry

You don't need a proprietary platform to start. OpenTelemetry (OTel) is a strong foundation for three reasons:

  • AI spans land in the same backend as your existing service traces, so you can see that a slow answer was really a slow database query inside a tool call.

  • It's vendor-neutral, so you can change backends without re-instrumenting.

  • It has emerging GenAI semantic conventions: standardized attribute names for models, token usage, and operations. They're still evolving, so pin your instrumentation library versions and expect some renaming over time.

Here's a minimal Python sketch that wraps an LLM call and a tool call:

from opentelemetry import trace

tracer = trace.get_tracer("support-agent")

def call_llm(client, messages, model, prompt_version):
    with tracer.start_as_current_span("llm_call") as span:
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("app.prompt_version", prompt_version)

        response = client.generate(model=model, messages=messages)

        span.set_attribute("gen_ai.usage.input_tokens", response.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", response.usage.output_tokens)
        span.set_attribute("gen_ai.response.finish_reason", response.finish_reason)
        return response


def call_tool(name, args, registry):
    with tracer.start_as_current_span(f"tool:{name}") as span:
        span.set_attribute("app.tool.name", name)
        span.set_attribute("app.tool.args_hash", hash_args(args))  # see redaction
        try:
            result = registry[name](**args)
            span.set_attribute("app.tool.success", True)
            return result
        except Exception as exc:
            span.record_exception(exc)
            span.set_attribute("app.tool.success", False)
            span.set_status(trace.Status(trace.StatusCode.ERROR))
            raise
from opentelemetry import trace

tracer = trace.get_tracer("support-agent")

def call_llm(client, messages, model, prompt_version):
    with tracer.start_as_current_span("llm_call") as span:
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("app.prompt_version", prompt_version)

        response = client.generate(model=model, messages=messages)

        span.set_attribute("gen_ai.usage.input_tokens", response.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", response.usage.output_tokens)
        span.set_attribute("gen_ai.response.finish_reason", response.finish_reason)
        return response


def call_tool(name, args, registry):
    with tracer.start_as_current_span(f"tool:{name}") as span:
        span.set_attribute("app.tool.name", name)
        span.set_attribute("app.tool.args_hash", hash_args(args))  # see redaction
        try:
            result = registry[name](**args)
            span.set_attribute("app.tool.success", True)
            return result
        except Exception as exc:
            span.record_exception(exc)
            span.set_attribute("app.tool.success", False)
            span.set_status(trace.Status(trace.StatusCode.ERROR))
            raise
from opentelemetry import trace

tracer = trace.get_tracer("support-agent")

def call_llm(client, messages, model, prompt_version):
    with tracer.start_as_current_span("llm_call") as span:
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("app.prompt_version", prompt_version)

        response = client.generate(model=model, messages=messages)

        span.set_attribute("gen_ai.usage.input_tokens", response.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", response.usage.output_tokens)
        span.set_attribute("gen_ai.response.finish_reason", response.finish_reason)
        return response


def call_tool(name, args, registry):
    with tracer.start_as_current_span(f"tool:{name}") as span:
        span.set_attribute("app.tool.name", name)
        span.set_attribute("app.tool.args_hash", hash_args(args))  # see redaction
        try:
            result = registry[name](**args)
            span.set_attribute("app.tool.success", True)
            return result
        except Exception as exc:
            span.record_exception(exc)
            span.set_attribute("app.tool.success", False)
            span.set_status(trace.Status(trace.StatusCode.ERROR))
            raise

And the agent loop that ties them together:

def run_agent(user_msg, max_steps=10):
    with tracer.start_as_current_span("agent_run") as run_span:
        run_span.set_attribute("app.max_steps", max_steps)
        messages = [{"role": "user", "content": user_msg}]

        for step in range(max_steps):
            reply = call_llm(client, messages, MODEL, PROMPT_VERSION)
            if not reply.tool_calls:
                run_span.set_attribute("app.steps_used", step + 1)
                run_span.set_attribute("app.outcome", "answered")
                return reply.text
            for tc in reply.tool_calls:
                messages.append(call_tool(tc.name, tc.args, TOOLS))

        run_span.set_attribute("app.outcome", "step_limit_hit")  # alert on this
        return fallback_response()
def run_agent(user_msg, max_steps=10):
    with tracer.start_as_current_span("agent_run") as run_span:
        run_span.set_attribute("app.max_steps", max_steps)
        messages = [{"role": "user", "content": user_msg}]

        for step in range(max_steps):
            reply = call_llm(client, messages, MODEL, PROMPT_VERSION)
            if not reply.tool_calls:
                run_span.set_attribute("app.steps_used", step + 1)
                run_span.set_attribute("app.outcome", "answered")
                return reply.text
            for tc in reply.tool_calls:
                messages.append(call_tool(tc.name, tc.args, TOOLS))

        run_span.set_attribute("app.outcome", "step_limit_hit")  # alert on this
        return fallback_response()
def run_agent(user_msg, max_steps=10):
    with tracer.start_as_current_span("agent_run") as run_span:
        run_span.set_attribute("app.max_steps", max_steps)
        messages = [{"role": "user", "content": user_msg}]

        for step in range(max_steps):
            reply = call_llm(client, messages, MODEL, PROMPT_VERSION)
            if not reply.tool_calls:
                run_span.set_attribute("app.steps_used", step + 1)
                run_span.set_attribute("app.outcome", "answered")
                return reply.text
            for tc in reply.tool_calls:
                messages.append(call_tool(tc.name, tc.args, TOOLS))

        run_span.set_attribute("app.outcome", "step_limit_hit")  # alert on this
        return fallback_response()

A few practical notes:

  • Wrap your LLM client once, in a single choke point, rather than sprinkling spans across the codebase. It keeps instrumentation consistent and easy to change.

  • Propagate context across service boundaries, including into tool services and queues, so the trace doesn't break in the middle.

  • Record the outcome explicitly. answered, step_limit_hit, escalated, and tool_failure are far more useful in queries than inferring status from logs.

  • Auto-instrumentation libraries exist for popular SDKs and agent frameworks. They're a fast way to start, but review what they capture, especially regarding sensitive content.

The metrics that matter

Beyond the usual request rate, error rate, and latency (RED metrics), agents need a few extra signals.

Cost and usage
  • Cost per request, computed from token counts and model pricing

  • Tokens per request (input vs. output), which is often where surprises hide

  • Cost per successful task, a better measure than cost per call, since failed and retried work still costs money

  • Cost per tenant or feature, so you can find the expensive outliers

Reliability
  • Tool-call failure rate, per tool

  • Retry counts and retry storms

  • Timeouts and rate-limit (429) responses from model providers

  • Provider fallback rate, if you fail over between models

Agent behavior
  • Steps per task, as a distribution (p50, p95, max)

  • Step-limit hits per hour

  • Tool-selection distribution, meaning which tools get called and how that shifts after a release

  • Refusal rate, since sudden changes often signal a prompt or model regression

Quality proxies

You can't measure "correctness" directly at scale, but you can watch signals that correlate with it:

  • User thumbs up/down ratio

  • Regeneration rate: how often users ask for another answer

  • Escalation-to-human rate

  • Abandonment after an agent response

  • Follow-up correction rate, such as users re-asking or rephrasing the same question

A sudden shift in any of these after a deploy is a strong hint that something changed, even when latency and errors look normal.

Failure modes worth alerting on

Agents fail in characteristic ways. Build alerts for these specifically:

Failure mode

What it looks like

Signal to alert on

Runaway loop

Agent repeats the same tool call or oscillates between steps

Steps per task above threshold; repeated identical tool calls

Cost blowout

One request or tenant burns huge token volume

Cost per request p99; tenant-level spend anomalies

Tool failure cascade

A downstream API degrades and the agent keeps retrying

Tool error rate by tool; retry counts

Silent regression

A prompt or model change degrades quality without errors

Quality proxies vs. baseline; eval scores on sampled traffic

Context overflow

Inputs exceed the window, so content is silently truncated

Input token counts near limit; truncation flags

Hallucinated tool args

Model invents IDs or parameters that don't exist

Validation errors, "not found" rates on tool calls

Prompt injection

Retrieved or user content hijacks the agent's behavior

Unexpected tool use; policy-violation classifier hits

The loop metric deserves special attention, because it's the failure that hurts your budget fastest. A stuck agent can make dozens of model calls in a minute. Pair your alert with a hard circuit breaker: a maximum number of steps, tokens, and wall-clock seconds per request, enforced in code. Alerts tell you about the problem, and limits stop it.

Handling sensitive data

Agent traces are uniquely dangerous to store. Prompts routinely contain personal data, credentials, medical or financial details, and proprietary documents, and tool arguments and results can be even more sensitive. Observability that creates a compliance incident isn't an improvement.

Decide these things deliberately, before you ship:

  • What gets stored raw? For many flows, metadata alone (token counts, latency, tool names, outcome) is enough.

  • Redaction: scrub known patterns (emails, phone numbers, API keys, card numbers) before data leaves your service, not afterward in the backend. Do it in a span processor or a collector pipeline.

  • Hashing: log a hash of tool arguments to detect repeats and loops without storing the values themselves.

  • Access control: restrict who can view full prompt and response content, and audit that access.

  • Retention: short by default. Keep raw content for days or weeks, aggregated metrics for longer.

  • Tenant isolation: ensure one customer's data can't surface in another's debug view.

  • Opt-in capture: make full-content capture a configuration switch you can turn on for a specific session while debugging, rather than an always-on default.

  • Regional and legal requirements: check where trace data is stored and whether deletion requests (for example, under privacy regulations) can be honored across your telemetry systems.

Involve your security and privacy teams early. It's much easier to design this in than to retrofit it.

Sampling without going blind

Storing every full trace with complete prompts is expensive, both in dollars and in risk. But naive random sampling throws away exactly the traces you most want.

Use tail-based sampling, which decides what to keep after the trace completes:

  • Keep 100% of traces with errors, tool failures, or step-limit hits

  • Keep 100% of slow requests (say, above p95 latency)

  • Keep 100% of traces with negative user feedback or human escalation

  • Keep 100% of traces from new prompt versions during a canary

  • Sample a small percentage (1 to 10%) of everything else as a baseline

The OpenTelemetry Collector supports tail-sampling policies for this. You retain the interesting cases at a fraction of the cost, and you always have a representative slice of normal traffic for comparison.

A useful separation: keep lightweight metrics for 100% of traffic (cheap, aggregated, no content) and rich traces for the interesting subset.

Measuring quality in production

Metrics and proxies tell you something changed. To know whether answers are actually good, add evaluation on live traffic:

  • Sample-based automated grading. Run an LLM judge (calibrated against human labels) over a random sample of production traces, scoring dimensions such as groundedness, correctness of tool use, and policy compliance. Track the scores as time series and alert on drops.

  • Deterministic checks on every response. Schema validity, presence of required citations, absence of forbidden content. They're cheap enough to run on everything.

  • Human review queues. Route low-confidence, negatively rated, or randomly sampled traces to reviewers. Give them a fast interface showing the full trace, not just the final answer.

  • Segment your results. Quality often differs by language, tenant, topic, or input length. An average can hide a serious problem in one segment.

  • Compare across versions. Tag every trace with prompt and model versions so you can compare a new release against the previous one on live data.

Remember that judges have biases and can drift, so periodically re-check them against human decisions.

Close the loop

The most valuable thing you can do with traces is feed them back into development.

   ┌────────────┐     ┌──────────────┐     ┌────────────────┐
   │ Production │────▶│ Trace + user │────▶│ Root-cause the │
   │  traffic   │     │   feedback   │     │    failure     │
   └────────────┘     └──────────────┘     └───────┬────────┘
         ▲                                         │
         │            ┌──────────────┐     ┌───────▼────────┐
         └────────────│ Ship the fix │◀────│ Add as eval /  │
                      │ (via CI/CD)  │     │ regression test│
                      └──────────────┘     └────────────────┘
   ┌────────────┐     ┌──────────────┐     ┌────────────────┐
   │ Production │────▶│ Trace + user │────▶│ Root-cause the │
   │  traffic   │     │   feedback   │     │    failure     │
   └────────────┘     └──────────────┘     └───────┬────────┘
         ▲                                         │
         │            ┌──────────────┐     ┌───────▼────────┐
         └────────────│ Ship the fix │◀────│ Add as eval /  │
                      │ (via CI/CD)  │     │ regression test│
                      └──────────────┘     └────────────────┘
   ┌────────────┐     ┌──────────────┐     ┌────────────────┐
   │ Production │────▶│ Trace + user │────▶│ Root-cause the │
   │  traffic   │     │   feedback   │     │    failure     │
   └────────────┘     └──────────────┘     └───────┬────────┘
         ▲                                         │
         │            ┌──────────────┐     ┌───────▼────────┐
         └────────────│ Ship the fix │◀────│ Add as eval /  │
                      │ (via CI/CD)  │     │ regression test│
                      └──────────────┘     └────────────────┘

When a user reports a bad answer, you should be able to:

  1. Find the exact trace from a session or request ID

  2. See where the reasoning went wrong: bad retrieval, wrong tool choice, hallucinated argument, misread tool output

  3. Fix the cause: a prompt change, a better tool description, a validation step, or a guardrail

  4. Turn the trace into a regression test in your eval set, with sensitive data scrubbed

  5. Ship the fix through your CI pipeline and watch the same metric in production

Production failures become tomorrow's test cases, and the system gets measurably more reliable with use. This loop is the real payoff of observability: it isn't a dashboard, it's a learning system.

Start simple: a rollout plan

You don't need a specialized platform on day one. A practical order of operations:

Week 1: Basic tracing

  • Wrap your LLM client and tool executor with spans

  • Record model, prompt version, token counts, latency, tool name, and outcome

  • View traces in whatever backend you already run

Week 2: Metrics and guardrails

  • Add cost per request, tool failure rate, and steps per task

  • Enforce hard limits on steps, tokens, and time

  • Alert on step-limit hits and cost outliers

Week 3: Data safety and sampling

  • Implement redaction and retention rules

  • Configure tail-based sampling to keep errors, slow, and unhappy-user traces

Week 4: Quality and feedback

  • Capture thumbs up/down and escalation events on the trace

  • Start a small sample-based judge, calibrated with human labels

  • Build the habit of turning bad traces into eval cases

Later, once you know which questions you keep asking, consider dedicated LLM observability tooling for prompt playgrounds, session replay, and dataset management. By then you'll know what you actually need.

Checklist
  • Every request is a trace; every LLM and tool call is a span

  • Model, prompt version, and token counts are recorded on LLM spans

  • Tool name, outcome, and errors are recorded on tool spans

  • Traces from different services are connected via context propagation

  • Cost per request and cost per successful task are tracked

  • Hard limits on steps, tokens, and wall-clock time are enforced in code

  • Alerts exist for loops, cost spikes, and tool failure cascades

  • Sensitive data is redacted before export, with defined retention and access rules

  • Tail-based sampling keeps every error, slow, and negatively rated trace

  • User feedback is attached to traces

  • A sampled judge (calibrated against humans) scores production quality

  • Bad traces routinely become eval cases

Closing thought

Agents don't fail like servers do. They fail like people: plausibly, confidently, and for reasons that only make sense once you see what they saw and what they tried. Observability is how you get to see that. Instrument the chain of decisions, watch cost and behavior alongside latency, protect sensitive data, and feed every failure back into your tests. The teams that do this find out about problems from their dashboards, not their customers.

Previous: The Real Cost of GPU Inference on Kubernetes · Start of series: Your Prompts Deserve a CI/CD Pipeline

2026 Elias, All rights reserved

2026 Elias, All rights reserved

Create a free website with Framer, the website builder loved by startups, designers and agencies.