Skip to content

Observability with OpenTelemetry

Export traces, metrics and log events from Agent SDK runs to any OTLP backend, link them to your own traces and tag them by agent and end user.

Once an agent runs in production you will want answers to questions like "which tool call made that run take four minutes?" and "which tenant burned the most tokens yesterday?". The Agent SDK can send OpenTelemetry traces, metrics and log events to any backend that speaks OTLP, whether that is a hosted platform or your own collector.

If you only need cost figures inside your app, read them from the message stream instead: see cost tracking.

How it works

The SDK launches the Claude Code CLI as a child process and talks to it over a local pipe. All the instrumentation lives in the CLI: spans around model requests and tool calls, counters for tokens and cost, structured events for prompts and tool results. The SDK itself emits nothing. It just passes configuration to the CLI, which exports straight to your collector.

Configuration is environment variables, which you can set in either place:

  • The process environment (shell, Dockerfile, Kubernetes manifest). Every query() picks them up with no code changes. This is what I use in production.
  • The env option per call, when different agents in one process need different settings. Python merges env with the inherited environment; TypeScript replaces it, so spread process.env first.

The three signals

Each signal has its own switch, so enable only what you need.

SignalContentsTurn on with
MetricsCounters for tokens, cost, sessions, lines of code, tool decisionsOTEL_METRICS_EXPORTER
Log eventsA record per prompt, API request, API error and tool resultOTEL_LOGS_EXPORTER
Traces (beta)Spans for interactions, model requests, tool calls and hooksOTEL_TRACES_EXPORTER and CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1

Metric names, event names and attributes are identical to the CLI's, documented in monitoring usage.

Turning export on

Nothing is exported until CLAUDE_CODE_ENABLE_TELEMETRY=1 is set and at least one exporter is chosen. A typical setup sends everything over OTLP/HTTP:

import { query } from "@anthropic-ai/claude-agent-sdk";

const telemetry = {
  CLAUDE_CODE_ENABLE_TELEMETRY: "1",
  CLAUDE_CODE_ENHANCED_TELEMETRY_BETA: "1",          // traces only
  OTEL_TRACES_EXPORTER: "otlp",
  OTEL_METRICS_EXPORTER: "otlp",
  OTEL_LOGS_EXPORTER: "otlp",
  OTEL_EXPORTER_OTLP_PROTOCOL: "http/protobuf",
  OTEL_EXPORTER_OTLP_ENDPOINT: "https://otel.internal.example:4318",
  OTEL_EXPORTER_OTLP_HEADERS: "x-api-key=REPLACE_ME"
};

for await (const msg of query({
  prompt: "Summarise yesterday's failed deploys from deploy.log",
  options: { env: { ...process.env, ...telemetry } }
})) {
  if (msg.type === "result") console.log(msg.subtype);
}

In Python, pass the same dictionary as ClaudeAgentOptions(env=telemetry); no spreading needed.

Warning: Never use the console exporter through the SDK. It writes to standard output, which is the channel the SDK uses for messages. To inspect telemetry locally, run an OpenTelemetry Collector on your machine and point OTEL_EXPORTER_OTLP_ENDPOINT at it.

Checking it works

Look in your collector's logs after a run. Export failures are silent by default: if the endpoint is down or rejects data, the agent carries on and the telemetry is dropped. To see exporter errors, set CLAUDE_CODE_OTEL_DIAG_STDERR=1 (Claude Code v2.1.179+) and read diagnostics through the SDK's stderr callback (Python) or stderr option (TypeScript).

Short-lived runs

The CLI batches telemetry and exports on a timer. It tries to flush on clean exit, but with a short timeout, so a slow collector can still lose spans, and a killed process loses whatever was buffered. Defaults are 60 seconds for metrics and 5 seconds for traces and logs. For short tasks, lower them:

VariableDefaultSuggested for short jobs
OTEL_METRIC_EXPORT_INTERVAL60000 ms1000
OTEL_LOGS_EXPORT_INTERVAL5000 ms1000
OTEL_TRACES_EXPORT_INTERVAL5000 ms1000

Reading traces

With the beta flag on, the agent loop becomes a tree of spans:

SpanCoversNotable attributes
claude_code.interactionOne turn, from prompt to responsesession.id
claude_code.llm_requestOne Claude API callModel, latency, token counts
claude_code.toolOne tool invocationChildren: claude_code.tool.blocked_on_user (waiting on permission) and claude_code.tool.execution
claude_code.hookOne hook runNeeds ENABLE_BETA_TRACING_DETAILED=1 and BETA_TRACING_ENDPOINT, which also changes where logs and traces go

llm_request, tool and hook spans sit under their interaction. When the agent delegates to a subagent, the subagent's llm_request and tool spans nest under the parent's claude_code.tool span, so one trace shows the whole delegation chain.

Spans carry session.id by default, so filtering on it stitches several query() calls in the same session into one timeline. Setting OTEL_METRICS_INCLUDE_SESSION_ID to a falsy value removes it.

Note: Tracing is beta. Span names and attributes can change between releases.

Token counts appear only when the API returned usage, so spans for failed or aborted requests may lack them.

Joining agent traces to your app's traces

The SDK propagates W3C trace context automatically. If an OpenTelemetry span is active in your code when you call query(), the SDK puts TRACEPARENT and TRACESTATE into the child environment and the CLI makes its claude_code.interaction span a child of yours. The agent run appears inside your HTTP request's trace rather than as a separate root.

Further behaviour:

  • OTLP log records from the run carry the same trace_id and span_id, so events join to spans. (Before v2.1.212, events emitted outside an active span lacked these IDs.)
  • The CLI forwards TRACEPARENT to every Bash and PowerShell command it runs. If one of those commands emits its own spans, they nest under the claude_code.tool.execution span.
  • If you set TRACEPARENT yourself in env, auto-injection is skipped, letting you pin a specific parent.
  • Interactive CLI sessions ignore an inbound TRACEPARENT. Only SDK and claude -p runs honour it.

Tagging by agent

Everything reports service.name as claude-code by default. With several agents sharing a collector, rename and tag each one:

opts = ClaudeAgentOptions(env={
    **telemetry,
    "OTEL_SERVICE_NAME": "invoice-reconciler",
    "OTEL_RESOURCE_ATTRIBUTES": "service.version=2.3.1,deployment.environment=staging",
})

These become resource attributes on every span, metric and event.

Attributing actions to end users

The CLI tags events with identity attributes derived from the credential it uses. In a multi-user service that is your service's key, not the person the agent was working for. Add the end user per call as resource attributes, percent-encoding values because OTEL_RESOURCE_ATTRIBUTES reserves commas, spaces and equals signs:

function telemetryFor(userId: string, orgId: string) {
  return {
    ...process.env,
    ...telemetry,
    OTEL_RESOURCE_ATTRIBUTES: `enduser.id=${encodeURIComponent(userId)},tenant.id=${encodeURIComponent(orgId)}`
  };
}

With that in place, the claude_code.tool_decision, claude_code.tool_result, claude_code.mcp_server_connection and claude_code.permission_mode_changed log records become a per-user audit trail you can ship to a SIEM. Monitoring usage lists every security-relevant event and its attributes.

What content gets exported

By default telemetry is structural: durations, model names, tool names and (when available) token counts. The content your agent reads and writes is not recorded. These opt-in variables add it:

VariableAdds
OTEL_LOG_USER_PROMPTS=1Prompt text on claude_code.user_prompt events and the claude_code.interaction span
OTEL_LOG_TOOL_DETAILS=1Tool arguments (file paths, shell commands, search patterns) on claude_code.tool_result events, plus real agent, skill, plugin and MCP server names on the cost and token metrics
OTEL_LOG_TOOL_CONTENT=1A tool.output span event on claude_code.tool with file contents, Bash output and MCP, WebFetch and WebSearch results. Needs tracing. Truncated at 60 KB by default; CLAUDE_CODE_OTEL_CONTENT_MAX_LENGTH changes that (v2.1.214+). MCP, WebFetch and WebSearch results need v2.1.283+.
OTEL_LOG_RAW_API_BODIESFull Messages API request and response JSON as claude_code.api_request_body and claude_code.api_response_body events. 1 inlines bodies (truncated at 60 KB by default, adjustable with CLAUDE_CODE_OTEL_CONTENT_MAX_LENGTH); file:<dir> writes untruncated bodies to disk with a body_ref path. Extended thinking is redacted. Enabling this exposes everything the other three would.

Span attributes carrying tool content have their own separate gates, described in monitoring usage.

Leave these off unless your telemetry pipeline is cleared to hold whatever data your agent handles. For client work I keep content logging off in production and only switch on OTEL_LOG_TOOL_DETAILS in staging.