Observability with OpenTelemetry
Export traces, metrics and log events from Agent SDK runs to any OTLP backend, link them to your own traces and tag them by agent and end user.
Once an agent runs in production you will want answers to questions like "which tool call made that run take four minutes?" and "which tenant burned the most tokens yesterday?". The Agent SDK can send OpenTelemetry traces, metrics and log events to any backend that speaks OTLP, whether that is a hosted platform or your own collector.
If you only need cost figures inside your app, read them from the message stream instead: see cost tracking.
How it works
The SDK launches the Claude Code CLI as a child process and talks to it over a local pipe. All the instrumentation lives in the CLI: spans around model requests and tool calls, counters for tokens and cost, structured events for prompts and tool results. The SDK itself emits nothing. It just passes configuration to the CLI, which exports straight to your collector.
Configuration is environment variables, which you can set in either place:
- The process environment (shell, Dockerfile, Kubernetes manifest). Every
query()picks them up with no code changes. This is what I use in production. - The
envoption per call, when different agents in one process need different settings. Python mergesenvwith the inherited environment; TypeScript replaces it, so spreadprocess.envfirst.
The three signals
Each signal has its own switch, so enable only what you need.
| Signal | Contents | Turn on with |
|---|---|---|
| Metrics | Counters for tokens, cost, sessions, lines of code, tool decisions | OTEL_METRICS_EXPORTER |
| Log events | A record per prompt, API request, API error and tool result | OTEL_LOGS_EXPORTER |
| Traces (beta) | Spans for interactions, model requests, tool calls and hooks | OTEL_TRACES_EXPORTER and CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 |
Metric names, event names and attributes are identical to the CLI's, documented in monitoring usage.
Turning export on
Nothing is exported until CLAUDE_CODE_ENABLE_TELEMETRY=1 is set and at least one exporter is chosen. A typical setup sends everything over OTLP/HTTP:
import { query } from "@anthropic-ai/claude-agent-sdk";
const telemetry = {
CLAUDE_CODE_ENABLE_TELEMETRY: "1",
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA: "1", // traces only
OTEL_TRACES_EXPORTER: "otlp",
OTEL_METRICS_EXPORTER: "otlp",
OTEL_LOGS_EXPORTER: "otlp",
OTEL_EXPORTER_OTLP_PROTOCOL: "http/protobuf",
OTEL_EXPORTER_OTLP_ENDPOINT: "https://otel.internal.example:4318",
OTEL_EXPORTER_OTLP_HEADERS: "x-api-key=REPLACE_ME"
};
for await (const msg of query({
prompt: "Summarise yesterday's failed deploys from deploy.log",
options: { env: { ...process.env, ...telemetry } }
})) {
if (msg.type === "result") console.log(msg.subtype);
}
In Python, pass the same dictionary as ClaudeAgentOptions(env=telemetry); no spreading needed.
Warning: Never use the
consoleexporter through the SDK. It writes to standard output, which is the channel the SDK uses for messages. To inspect telemetry locally, run an OpenTelemetry Collector on your machine and pointOTEL_EXPORTER_OTLP_ENDPOINTat it.
Checking it works
Look in your collector's logs after a run. Export failures are silent by default: if the endpoint is down or rejects data, the agent carries on and the telemetry is dropped. To see exporter errors, set CLAUDE_CODE_OTEL_DIAG_STDERR=1 (Claude Code v2.1.179+) and read diagnostics through the SDK's stderr callback (Python) or stderr option (TypeScript).
Short-lived runs
The CLI batches telemetry and exports on a timer. It tries to flush on clean exit, but with a short timeout, so a slow collector can still lose spans, and a killed process loses whatever was buffered. Defaults are 60 seconds for metrics and 5 seconds for traces and logs. For short tasks, lower them:
| Variable | Default | Suggested for short jobs |
|---|---|---|
OTEL_METRIC_EXPORT_INTERVAL | 60000 ms | 1000 |
OTEL_LOGS_EXPORT_INTERVAL | 5000 ms | 1000 |
OTEL_TRACES_EXPORT_INTERVAL | 5000 ms | 1000 |
Reading traces
With the beta flag on, the agent loop becomes a tree of spans:
| Span | Covers | Notable attributes |
|---|---|---|
claude_code.interaction | One turn, from prompt to response | session.id |
claude_code.llm_request | One Claude API call | Model, latency, token counts |
claude_code.tool | One tool invocation | Children: claude_code.tool.blocked_on_user (waiting on permission) and claude_code.tool.execution |
claude_code.hook | One hook run | Needs ENABLE_BETA_TRACING_DETAILED=1 and BETA_TRACING_ENDPOINT, which also changes where logs and traces go |
llm_request, tool and hook spans sit under their interaction. When the agent delegates to a subagent, the subagent's llm_request and tool spans nest under the parent's claude_code.tool span, so one trace shows the whole delegation chain.
Spans carry session.id by default, so filtering on it stitches several query() calls in the same session into one timeline. Setting OTEL_METRICS_INCLUDE_SESSION_ID to a falsy value removes it.
Note: Tracing is beta. Span names and attributes can change between releases.
Token counts appear only when the API returned usage, so spans for failed or aborted requests may lack them.
Joining agent traces to your app's traces
The SDK propagates W3C trace context automatically. If an OpenTelemetry span is active in your code when you call query(), the SDK puts TRACEPARENT and TRACESTATE into the child environment and the CLI makes its claude_code.interaction span a child of yours. The agent run appears inside your HTTP request's trace rather than as a separate root.
Further behaviour:
- OTLP log records from the run carry the same
trace_idandspan_id, so events join to spans. (Before v2.1.212, events emitted outside an active span lacked these IDs.) - The CLI forwards
TRACEPARENTto every Bash and PowerShell command it runs. If one of those commands emits its own spans, they nest under theclaude_code.tool.executionspan. - If you set
TRACEPARENTyourself inenv, auto-injection is skipped, letting you pin a specific parent. - Interactive CLI sessions ignore an inbound
TRACEPARENT. Only SDK andclaude -pruns honour it.
Tagging by agent
Everything reports service.name as claude-code by default. With several agents sharing a collector, rename and tag each one:
opts = ClaudeAgentOptions(env={
**telemetry,
"OTEL_SERVICE_NAME": "invoice-reconciler",
"OTEL_RESOURCE_ATTRIBUTES": "service.version=2.3.1,deployment.environment=staging",
})
These become resource attributes on every span, metric and event.
Attributing actions to end users
The CLI tags events with identity attributes derived from the credential it uses. In a multi-user service that is your service's key, not the person the agent was working for. Add the end user per call as resource attributes, percent-encoding values because OTEL_RESOURCE_ATTRIBUTES reserves commas, spaces and equals signs:
function telemetryFor(userId: string, orgId: string) {
return {
...process.env,
...telemetry,
OTEL_RESOURCE_ATTRIBUTES: `enduser.id=${encodeURIComponent(userId)},tenant.id=${encodeURIComponent(orgId)}`
};
}
With that in place, the claude_code.tool_decision, claude_code.tool_result, claude_code.mcp_server_connection and claude_code.permission_mode_changed log records become a per-user audit trail you can ship to a SIEM. Monitoring usage lists every security-relevant event and its attributes.
What content gets exported
By default telemetry is structural: durations, model names, tool names and (when available) token counts. The content your agent reads and writes is not recorded. These opt-in variables add it:
| Variable | Adds |
|---|---|
OTEL_LOG_USER_PROMPTS=1 | Prompt text on claude_code.user_prompt events and the claude_code.interaction span |
OTEL_LOG_TOOL_DETAILS=1 | Tool arguments (file paths, shell commands, search patterns) on claude_code.tool_result events, plus real agent, skill, plugin and MCP server names on the cost and token metrics |
OTEL_LOG_TOOL_CONTENT=1 | A tool.output span event on claude_code.tool with file contents, Bash output and MCP, WebFetch and WebSearch results. Needs tracing. Truncated at 60 KB by default; CLAUDE_CODE_OTEL_CONTENT_MAX_LENGTH changes that (v2.1.214+). MCP, WebFetch and WebSearch results need v2.1.283+. |
OTEL_LOG_RAW_API_BODIES | Full Messages API request and response JSON as claude_code.api_request_body and claude_code.api_response_body events. 1 inlines bodies (truncated at 60 KB by default, adjustable with CLAUDE_CODE_OTEL_CONTENT_MAX_LENGTH); file:<dir> writes untruncated bodies to disk with a body_ref path. Extended thinking is redacted. Enabling this exposes everything the other three would. |
Span attributes carrying tool content have their own separate gates, described in monitoring usage.
Leave these off unless your telemetry pipeline is cleared to hold whatever data your agent handles. For client work I keep content logging off in production and only switch on OTEL_LOG_TOOL_DETAILS in staging.