Skip to content

Cost tracking

Read token usage and estimated cost from Agent SDK messages, avoid double counting, handle streaming and crashes, and tune prompt cache TTLs.

Every Agent SDK run reports what it used: tokens per step, tokens per model and an estimated dollar cost at the end. That is enough to build per-customer usage dashboards, budget alerts and "this run cost about 4p" footers. This page explains where each number lives, what it includes, and the handful of traps that produce wrong totals.

Warning: total_cost_usd and costUSD are estimates computed on the client from a price table bundled with the SDK (or from a modelPricing table if one is set). They drift when prices change, when your SDK version does not know a model, or when billing rules apply that the client cannot see. Use them for insight and rough budgets. For real billing, use the Usage and Cost API or the Usage page in the Claude Console, and never invoice customers from these fields.

One billing rule the SDK does model is data residency: when a response's usage reports inference_geo: "us", its token cost is multiplied by 1.1 (per-request fees such as web search are not). This needs TypeScript SDK v0.3.239 or Python SDK v0.2.144.

Three scopes

ScopeWhat it isWhere cost appears
StepOne request and response with the modelusage on assistant messages
query() callOne call to query(), possibly many stepsThe result message at the end (one per turn in streaming input mode)
SessionCalls linked through resumeEach resumed call's result includes earlier spend

Field names in each SDK

DataTypeScriptPython
Per-step usagemessage.message.usagemessage.usage
Step IDmessage.message.idmessage.message_id
Per-model breakdownresult.modelUsageresult.model_usage
Total estimateresult.total_cost_usdresult.total_cost_usd (optional, may be None)

The model and granularity are identical; only names and nesting differ.

The total for a call

The result message ends a query() call and carries total_cost_usd. Both success and error results include it, so read it whatever the subtype.

from claude_agent_sdk import query, ResultMessage

try:
    async for msg in query(prompt="Write a migration plan for moving uploads to S3"):
        if isinstance(msg, ResultMessage):
            print(f"Estimated cost: ${msg.total_cost_usd or 0:.4f} ({msg.subtype})")
except Exception as err:
    print("Run failed:", err)   # the result above has already been printed

What each result field counts

This matters as soon as subagents are involved:

FieldIncludes subagents?
usageNo. Top-level loop only.
total_cost_usdYes
modelUsage / model_usageYes, split by model

So for whole-run token accounting, use modelUsage. usage undercounts the moment anything is delegated.

In single message input mode, if background subagents are still running when the final turn ends, Claude Code waits for them (up to the cap described in headless mode) before emitting the result, and the totals include that work. To keep subagent spend bounded in the first place, use the depth, concurrency and budget limits.

Per-step usage without double counting

When Claude calls several tools in parallel, you get several assistant messages that share one message ID and identical usage. Count each ID once.

import { query } from "@anthropic-ai/claude-agent-sdk";

const counted = new Set<string>();
let input = 0, cacheRead = 0, output = 0;

for await (const msg of query({ prompt: "Audit our npm dependencies for licence issues" })) {
  if (msg.type === "assistant" && !msg.parent_tool_use_id) {
    const { id, usage } = msg.message;
    if (!counted.has(id)) {
      counted.add(id);
      input += usage.input_tokens;
      cacheRead += usage.cache_read_input_tokens ?? 0;
    }
  }
  if (msg.type === "result") output = msg.usage.output_tokens;
}
console.log({ steps: counted.size, input, cacheRead, output });

Skipping messages with a parent_tool_use_id keeps subagent traffic out of the main-loop count.

Output tokens: read them from the result

Per-step output_tokens is a placeholder. Claude Code builds each assistant message from the usage reported at message_start, before any output was generated, and every message from the same response repeats that number. The true output count arrives at the end of the response and is added to the result message. Read output tokens from the result's usage, or from modelUsage.

To watch output grow live, turn on includePartialMessages (include_partial_messages) and read usage from each message_delta stream event.

Per-model breakdown

modelUsage maps model name to counts and cost. It is the quickest way to see whether a Haiku subagent is actually saving you money compared with the Opus main loop.

for await (const msg of run) {
  if (msg.type !== "result") continue;
  for (const [model, u] of Object.entries(msg.modelUsage)) {
    console.log(model, `$${u.costUSD.toFixed(4)}`, {
      in: u.inputTokens, out: u.outputTokens,
      cacheRead: u.cacheReadInputTokens, cacheWrite: u.cacheCreationInputTokens,
      basis: u.costBasis
    });
  }
}

costBasis (Claude Code v2.1.246+) tells you which table priced the model's latest request: list, managed (a modelPricing table) or unknown (no match for the model ID). An unknown is your cue to upgrade the SDK or add pricing.

Totals across several calls

  • Independent calls (no resume or continue): each result covers only itself. Add them up.
  • Calls resuming one session: Claude Code stores the session's totals in its transcript when the process exits normally and restores them on resume or fork. Each result already includes earlier spend, so read the latest one; summing double counts. Before v2.1.277, a resumed session started from zero.
grand_total = 0.0
for task in ["Summarise src/billing", "List every public function in src/billing/api.py"]:
    async for msg in query(prompt=task):
        if isinstance(msg, ResultMessage):
            grand_total += msg.total_cost_usd or 0
print(f"Batch total: ${grand_total:.4f}")

Streaming input mode

In streaming input mode one query() call carries many user turns and each turn emits a result.

FieldScope
usageThat turn only, main loop only
total_cost_usd, modelUsage / model_usageRunning total for the whole call, plus anything restored from a resumed session

If your app never sends /clear, /reset or /new, simply read the latest result.

Those three commands reset the running totals (nothing else does inside a call):

  • The reset turn's own result counts only from the reset and has a new session_id.
  • Later results keep counting from that reset.
  • The last result before each reset holds the total since the previous reset.

So the call's total is the sum of the last result before each reset, plus the final result. TypeScript emits an SDKConversationResetMessage at each reset; Python emits ConversationResetMessage (dropped by the iterator before Python SDK v0.2.137, so count /clear turns yourself on older versions).

maxBudgetUsd / max_budget_usd counts only the call's own spend: restored session totals do not count against it, and a /clear restarts the budget.

Failed runs

Failures still cost money up to the point of failure. Read cost from every result, whatever the subtype, and note two cases where usage under-reports:

ResultCaveat
error_during_execution after a crashEvery cost field may be zero
error_max_budget_usdusage omits the response that crossed the budget; total_cost_usd and modelUsage include it

Prefer total_cost_usd or modelUsage for accounting.

Recovering after a crash

When the Claude Code process crashes it emits a final error_during_execution result that may be zeroed, in both single-shot and streaming modes.

  1. In streaming mode, use the result from the turn before the crash; it holds the running total. This does not work for single-shot calls, crashes on the first turn, or when the previous turn was the /clear itself.
  2. Otherwise, sum usage across the assistant messages (deduplicated by ID): all of them for single-shot, or those after the last result for streaming. This recovers main-loop input and cache tokens only. Subagent usage, output tokens and dollar cost cannot be recovered this way.

Prompt caching

Caching is automatic in the SDK. Usage objects carry two extra fields:

  • cache_creation_input_tokens: tokens written to cache, charged above normal input.
  • cache_read_input_tokens: tokens read from cache, charged well below normal input.

In TypeScript they are typed on Usage. In Python they are dictionary keys, for example msg.usage.get("cache_read_input_tokens", 0). A healthy long-running agent shows reads dwarfing writes; if not, look at whether your system prompt varies between runs (see modifying system prompts).

Longer cache lifetimes

Requests fall into two TTL buckets: your own conversation turns (plus helpers Claude Code runs inline with them) and everything else, such as subagents.

With an API key, or on Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS, your turns default to a 5-minute TTL. If you run many short sessions against the same prompt with gaps longer than five minutes, the cache expires between them and you pay full input price each time.

ControlApplies toValues
ENABLE_PROMPT_CACHING_1HBoth bucketsSet to request 1-hour writes everywhere
CLAUDE_CODE_PROMPT_CACHE_TTL or the promptCacheTtl settingMain conversation5m or 1h
CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL or the subagentPromptCacheTtl settingEverything else5m or 1h

The per-bucket controls override ENABLE_PROMPT_CACHING_1H. One-hour writes cost more than five-minute writes, so you are trading write cost for more reads.

const options = {
  env: { ...process.env, CLAUDE_CODE_USE_BEDROCK: "1", ENABLE_PROMPT_CACHING_1H: "1" }
};

That example needs working AWS credentials for Amazon Bedrock. On a Claude subscription within included usage you already get the 1-hour TTL on your own turns (and some helpers); Claude Code drops to 5 minutes once you are on usage credits, unless promptCacheTtl is 1h. The prompt caching page has the full precedence order.