Cost tracking
Read token usage and estimated cost from Agent SDK messages, avoid double counting, handle streaming and crashes, and tune prompt cache TTLs.
Every Agent SDK run reports what it used: tokens per step, tokens per model and an estimated dollar cost at the end. That is enough to build per-customer usage dashboards, budget alerts and "this run cost about 4p" footers. This page explains where each number lives, what it includes, and the handful of traps that produce wrong totals.
Warning:
total_cost_usdandcostUSDare estimates computed on the client from a price table bundled with the SDK (or from amodelPricingtable if one is set). They drift when prices change, when your SDK version does not know a model, or when billing rules apply that the client cannot see. Use them for insight and rough budgets. For real billing, use the Usage and Cost API or the Usage page in the Claude Console, and never invoice customers from these fields.
One billing rule the SDK does model is data residency: when a response's usage reports inference_geo: "us", its token cost is multiplied by 1.1 (per-request fees such as web search are not). This needs TypeScript SDK v0.3.239 or Python SDK v0.2.144.
Three scopes
| Scope | What it is | Where cost appears |
|---|---|---|
| Step | One request and response with the model | usage on assistant messages |
query() call | One call to query(), possibly many steps | The result message at the end (one per turn in streaming input mode) |
| Session | Calls linked through resume | Each resumed call's result includes earlier spend |
Field names in each SDK
| Data | TypeScript | Python |
|---|---|---|
| Per-step usage | message.message.usage | message.usage |
| Step ID | message.message.id | message.message_id |
| Per-model breakdown | result.modelUsage | result.model_usage |
| Total estimate | result.total_cost_usd | result.total_cost_usd (optional, may be None) |
The model and granularity are identical; only names and nesting differ.
The total for a call
The result message ends a query() call and carries total_cost_usd. Both success and error results include it, so read it whatever the subtype.
from claude_agent_sdk import query, ResultMessage
try:
async for msg in query(prompt="Write a migration plan for moving uploads to S3"):
if isinstance(msg, ResultMessage):
print(f"Estimated cost: ${msg.total_cost_usd or 0:.4f} ({msg.subtype})")
except Exception as err:
print("Run failed:", err) # the result above has already been printed
What each result field counts
This matters as soon as subagents are involved:
| Field | Includes subagents? |
|---|---|
usage | No. Top-level loop only. |
total_cost_usd | Yes |
modelUsage / model_usage | Yes, split by model |
So for whole-run token accounting, use modelUsage. usage undercounts the moment anything is delegated.
In single message input mode, if background subagents are still running when the final turn ends, Claude Code waits for them (up to the cap described in headless mode) before emitting the result, and the totals include that work. To keep subagent spend bounded in the first place, use the depth, concurrency and budget limits.
Per-step usage without double counting
When Claude calls several tools in parallel, you get several assistant messages that share one message ID and identical usage. Count each ID once.
import { query } from "@anthropic-ai/claude-agent-sdk";
const counted = new Set<string>();
let input = 0, cacheRead = 0, output = 0;
for await (const msg of query({ prompt: "Audit our npm dependencies for licence issues" })) {
if (msg.type === "assistant" && !msg.parent_tool_use_id) {
const { id, usage } = msg.message;
if (!counted.has(id)) {
counted.add(id);
input += usage.input_tokens;
cacheRead += usage.cache_read_input_tokens ?? 0;
}
}
if (msg.type === "result") output = msg.usage.output_tokens;
}
console.log({ steps: counted.size, input, cacheRead, output });
Skipping messages with a parent_tool_use_id keeps subagent traffic out of the main-loop count.
Output tokens: read them from the result
Per-step output_tokens is a placeholder. Claude Code builds each assistant message from the usage reported at message_start, before any output was generated, and every message from the same response repeats that number. The true output count arrives at the end of the response and is added to the result message. Read output tokens from the result's usage, or from modelUsage.
To watch output grow live, turn on includePartialMessages (include_partial_messages) and read usage from each message_delta stream event.
Per-model breakdown
modelUsage maps model name to counts and cost. It is the quickest way to see whether a Haiku subagent is actually saving you money compared with the Opus main loop.
for await (const msg of run) {
if (msg.type !== "result") continue;
for (const [model, u] of Object.entries(msg.modelUsage)) {
console.log(model, `$${u.costUSD.toFixed(4)}`, {
in: u.inputTokens, out: u.outputTokens,
cacheRead: u.cacheReadInputTokens, cacheWrite: u.cacheCreationInputTokens,
basis: u.costBasis
});
}
}
costBasis (Claude Code v2.1.246+) tells you which table priced the model's latest request: list, managed (a modelPricing table) or unknown (no match for the model ID). An unknown is your cue to upgrade the SDK or add pricing.
Totals across several calls
- Independent calls (no
resumeorcontinue): each result covers only itself. Add them up. - Calls resuming one session: Claude Code stores the session's totals in its transcript when the process exits normally and restores them on resume or fork. Each result already includes earlier spend, so read the latest one; summing double counts. Before v2.1.277, a resumed session started from zero.
grand_total = 0.0
for task in ["Summarise src/billing", "List every public function in src/billing/api.py"]:
async for msg in query(prompt=task):
if isinstance(msg, ResultMessage):
grand_total += msg.total_cost_usd or 0
print(f"Batch total: ${grand_total:.4f}")
Streaming input mode
In streaming input mode one query() call carries many user turns and each turn emits a result.
| Field | Scope |
|---|---|
usage | That turn only, main loop only |
total_cost_usd, modelUsage / model_usage | Running total for the whole call, plus anything restored from a resumed session |
If your app never sends /clear, /reset or /new, simply read the latest result.
Those three commands reset the running totals (nothing else does inside a call):
- The reset turn's own result counts only from the reset and has a new
session_id. - Later results keep counting from that reset.
- The last result before each reset holds the total since the previous reset.
So the call's total is the sum of the last result before each reset, plus the final result. TypeScript emits an SDKConversationResetMessage at each reset; Python emits ConversationResetMessage (dropped by the iterator before Python SDK v0.2.137, so count /clear turns yourself on older versions).
maxBudgetUsd / max_budget_usd counts only the call's own spend: restored session totals do not count against it, and a /clear restarts the budget.
Failed runs
Failures still cost money up to the point of failure. Read cost from every result, whatever the subtype, and note two cases where usage under-reports:
| Result | Caveat |
|---|---|
error_during_execution after a crash | Every cost field may be zero |
error_max_budget_usd | usage omits the response that crossed the budget; total_cost_usd and modelUsage include it |
Prefer total_cost_usd or modelUsage for accounting.
Recovering after a crash
When the Claude Code process crashes it emits a final error_during_execution result that may be zeroed, in both single-shot and streaming modes.
- In streaming mode, use the result from the turn before the crash; it holds the running total. This does not work for single-shot calls, crashes on the first turn, or when the previous turn was the
/clearitself. - Otherwise, sum
usageacross the assistant messages (deduplicated by ID): all of them for single-shot, or those after the last result for streaming. This recovers main-loop input and cache tokens only. Subagent usage, output tokens and dollar cost cannot be recovered this way.
Prompt caching
Caching is automatic in the SDK. Usage objects carry two extra fields:
cache_creation_input_tokens: tokens written to cache, charged above normal input.cache_read_input_tokens: tokens read from cache, charged well below normal input.
In TypeScript they are typed on Usage. In Python they are dictionary keys, for example msg.usage.get("cache_read_input_tokens", 0). A healthy long-running agent shows reads dwarfing writes; if not, look at whether your system prompt varies between runs (see modifying system prompts).
Longer cache lifetimes
Requests fall into two TTL buckets: your own conversation turns (plus helpers Claude Code runs inline with them) and everything else, such as subagents.
With an API key, or on Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS, your turns default to a 5-minute TTL. If you run many short sessions against the same prompt with gaps longer than five minutes, the cache expires between them and you pay full input price each time.
| Control | Applies to | Values |
|---|---|---|
ENABLE_PROMPT_CACHING_1H | Both buckets | Set to request 1-hour writes everywhere |
CLAUDE_CODE_PROMPT_CACHE_TTL or the promptCacheTtl setting | Main conversation | 5m or 1h |
CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL or the subagentPromptCacheTtl setting | Everything else | 5m or 1h |
The per-bucket controls override ENABLE_PROMPT_CACHING_1H. One-hour writes cost more than five-minute writes, so you are trading write cost for more reads.
const options = {
env: { ...process.env, CLAUDE_CODE_USE_BEDROCK: "1", ENABLE_PROMPT_CACHING_1H: "1" }
};
That example needs working AWS credentials for Amazon Bedrock. On a Claude subscription within included usage you already get the 1-hour TTL on your own turns (and some helpers); Claude Code drops to 5 minutes once you are on usage credits, unless promptCacheTtl is 1h. The prompt caching page has the full precedence order.