Back to News & insightsEngineering

AI agent observability: trace the work without recording everything

Connect model calls, tools, retries, and validation outcomes in a useful trace while minimizing sensitive content and misleading success metrics.

Editorial guide · Updated September 27, 2026 · 4 min read
Branching rails of glass event beads reveal a blocked connection and an alternate path.

An agent says the task is complete. The user finds an empty report. Without a record of the work, the team cannot tell whether retrieval failed, a tool timed out, or the model described an action it never performed.

Observability should connect the requested outcome with the events that produced it. That does not require storing every private prompt and document indefinitely. It requires deliberate event boundaries, safe identifiers, and a definition of success outside the assistant's final sentence.

Trace a task, not just a model request

OpenTelemetry describes traces as related spans representing work across an execution path. That model is useful for an agent because one user request can trigger retrieval, several model calls, tool execution, validation, and retries.

Give the top-level task a stable identifier and connect its operations. Record safe metadata such as component version, duration, status, and a bounded failure category. Preserve relationships between retries so they do not appear to be independent successful tasks.

An example of a misleading success

Imagine a fictional research agent preparing a bibliography. Search succeeds, one source download fails, and the final response claims that all sources were checked. The model call itself returned successfully, but the user-level task did not meet its contract.

Define the contract before instrumenting: each included source must have a verified identifier, supported bibliographic fields, and an accessible provenance record. The trace should capture the validation outcome, not merely the HTTP status of the model service.

Separate transport success, tool success, and task acceptance. This makes it possible to see a system whose infrastructure is healthy while its completed work is unreliable.

Record decisions at consequential boundaries

For the bibliography agent, useful events include retrieval completion, source-fetch outcome, citation validation, and final artifact creation. A tool dispatch should record whether authorization passed and which safe operation identity was used.

Avoid logging full document bodies simply because they are available. Often a source identifier, version, byte count, and validation code are enough to locate a problem. If deeper inspection is needed, use a controlled diagnostic path with appropriate access and retention.

The OpenTelemetry security documentation highlights the need to protect sensitive telemetry. Treat the observability store as part of the application's data boundary, not as an unrestricted dumping ground for everything the model saw.

Make latency explainable

Break total waiting time into meaningful stages: queueing, retrieval, model generation, tool execution, and final validation. If the bibliography agent spends most of its time downloading sources, switching language models may barely change the user experience.

Distinguish parallel and sequential work. Adding every span duration can overstate elapsed time when operations overlap. Inspect the critical path that actually delays completion.

Track cancellations and abandoned tasks. A user who closes the page after a long wait should not disappear from the reliability picture. Record whether underlying work stopped, completed later, or required cleanup.

Sample with a purpose

High-volume systems may not retain every detailed trace. Choose a sampling policy that preserves enough information about failures, slow tasks, and important workflow branches. Keep aggregate counters so sampling does not make the denominator ambiguous.

For rare errors, a trace chosen only at request start may miss the interesting outcome. Consider supported sampling strategies that can retain failed or slow executions while respecting data and resource limits. Document the policy so analysts know what the trace collection represents.

Do not infer an overall failure rate directly from a trace set intentionally enriched for failures. Use the appropriate aggregate measurement and use detailed traces to explain the cases.

Turn traces into better evaluations

Review recurring failure patterns and create sanitized test cases. For example, a source with a missing publication year can become a fixture that verifies the agent leaves the field unresolved instead of inventing a date.

Keep those fixtures independent of private production content where possible. Record the behavior being tested and the expected outcome. The value of an incident is the generalizable lesson, not the indefinite retention of the original user's material.

Connect a release to the component versions visible in traces. When behavior changes, the team should be able to identify which prompt, model, retrieval configuration, or validator was active.

Make completion an observable fact

A reliable completion event should correspond to an accepted artifact or verified state change. If validation fails, preserve that distinction in the user interface and the metrics. A confident final message must not overwrite the operational record.

Good agent observability explains where work went wrong and whether users received what they requested. It does so with enough evidence to investigate, and enough restraint to avoid creating an unnecessary archive of everything people entrusted to the assistant.

Technical background

Read source on opentelemetry.io

Read source on opentelemetry.io

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.