Skip to content

Manage costs

Track Claude Code spend with /usage and /insights, control it across Teams, Enterprise, Console and cloud setups, and cut token use day to day.

Claude Code is metered in tokens. Subscription plans (Pro, Max, Team, Enterprise) wrap that in a plan allowance; API and cloud-provider access bills each token. Either way, what you spend depends mostly on the model you pick, how big your context gets, and whether you run several sessions or automation in parallel.

For a planning figure: across enterprise deployments the average is about $13 per developer per active day, or $150 to $250 per developer per month, and 90% of users stay under $30 per active day. I always tell clients to run a small pilot first and use the tools below to get their own baseline before committing to a number.

This page covers Claude Code only. Usage limits for other Claude products are documented in Anthropic's Help Centre.

Track your own usage

/usage

The Session block at the top of /usage breaks down tokens for the current session:

Total cost:            $1.92
Total duration (API):  14m 05s
Total duration (wall): 2h 41m 37s
Total code changes:    312 lines added, 88 lines removed
Usage by model:
   claude-opus-4-8:  3.4k input, 18.9k output, 2.1m cache read, 140.0k cache write ($1.81)
   claude-haiku-4-5:  9.2k input, 1.1k output ($0.11)

Things to know about that figure:

  • It is computed on your machine from token counts at list price, so it is an estimate. The Usage page in the Claude Console is the authoritative bill.
  • If an admin has set a modelPricing table in managed settings, the figure uses your contracted rates and the Total cost line says at your organization's configured rates (see below).
  • Responses billed at the 1.1x data residency rate are multiplied by 1.1. The same figure feeds the status line cost field and counts towards --max-budget-usd (see CLI reference).
  • /clear starts a new session, so totals reset to zero. Before v2.1.211 they accumulated for the life of the process.
  • For Pro and Max subscribers the dollar figure is not what you pay; your usage is in the plan. You also see plan usage bars, activity stats and a breakdown on the same screen.

Prompt cache line

From v2.1.251, after the first response a Prompt cache (main) line summarises how well prompt caching is working: request count, share of input tokens served from cache, misses, and whether the cache is currently warm. For example:

Prompt cache (main):   22 requests · 88% of input tokens from cache · 3 misses (last 12m ago, 205.4k tokens re-cached, likely cause: tool definitions changed) · warm (5m TTL, last activity 1m ago)
  • Misses are requests that reprocessed content the cache already held, with when the last one happened and how many tokens were re-written. From v2.1.260 a likely cause is named when Claude Code can tell.
  • Expected rebuilds appear once compaction or tool-result clearing has rewritten the conversation; those misses are counted separately because Claude Code caused them on purpose.
  • Warm or cold says whether the cached prefix is still within its lifetime and which TTL applies. Cold shows idle time. If no response reported cache tokens, the line ends no prompt caching reported by the API.

It reads the cache fields in API responses, so it works on any provider or gateway. It covers the main conversation, not subagents, and resets with /clear. Status line scripts get the same numbers from the prompt_cache object.

Plan usage breakdown

On Pro, Max, Team and Enterprise, /usage also explains what is eating your allowance:

  • Attribution: percentage of recent usage from skills, subagents, plugins and individual MCP servers. An MCP server is only charged for requests that consumed one of its tool results (before v2.1.222 it was over-counted).
  • Behaviour flags: anything such as long context or cache misses that accounts for 10% or more of recent usage, each with a tip.
  • Loops (v2.1.242+): the heaviest /loop and other scheduled tasks, with how often each fires, run count, total and per-run tokens, and last run. Rows are keyed on the prompt, so stopping and recreating a loop keeps one row.

Press d or w to flip between 24 hours and 7 days. Figures are approximate and come from local history, so other devices and claude.ai are excluded. In VS Code, the attribution and flags appear in the Account & usage dialog with a Day/Week toggle, without Loops.

Usage-credits row

While usage credits are on, /usage shows a row for them:

  • Pro and Max: this month's spend against your monthly limit, or Unlimited with no figure if you have not set one.
  • Team and Enterprise: your own spend this month against any limit that applies to you personally. Organisation-wide limits are not shown here. No personal limit means spend with nothing beside it. If credits are off for you, there is no row.

With a limit set, the row appears straight away at 0% (before v2.1.236 it only appeared on Pro and Max, and only after first spend).

If the usage request fails

Usually because the usage endpoint is rate limited. /usage then shows the last bars it loaded on this machine in the past 60 minutes with a Showing last-known usage note and the data's age. Press r to retry. With no recent snapshot it reports the rate limit and offers the same retry.

/insights

/insights is about how you work, not how much. It analyses recent sessions on this machine and writes an HTML report on what you work on, where things went wrong (misunderstood requests, buggy code) and how to get more out of Claude Code. Each run takes up to 200 sessions it has not seen before and skips very short ones; if some are left out, the header shows something like 200 sessions (412 total). If auto mode is available and you mostly worked without it, the report can estimate how many permission prompts it would have handled.

The latest report is ~/.claude/usage-data/report.html, with a timestamped copy per run alongside. Reports are cleaned up with other session data after cleanupPeriodDays (30 days by default). It works on any plan and provider, uses your normal account so it costs tokens, and ignores other devices and claude.ai. See Commands.

Usage credits on a subscription

Usage credits let you keep going past your plan limit. Run /usage-credits while signed in with a claude.ai subscription (it does not work with API key auth). Self-serve Enterprise, Enterprise trials and Enterprise billed through AWS Marketplace need v2.1.248 or later; older versions report Unknown command: /usage-credits.

You are/usage-credits does
A Pro or Max subscriberOpens Settings, Usage on claude.ai, where you switch credits on or off and see balance, monthly spend and limit
Team or Enterprise with billing accessOpens Organisation settings, Usage
Team or Enterprise without billing accessAsks you to confirm, then sends a request to your admins

The request flow only works interactively; with -p or from Remote Control it tells you to use an interactive session. Repeating it while a request is pending tells you one is already waiting; once an admin dismisses it, you can send another (before v2.1.222 a dismissal blocked new requests). On Pro and Max, hitting your spend limit with credits left prompts you to raise or remove the limit without leaving the CLI.

Manage costs for an organisation

Your controls depend on how people sign in. Teams and Enterprise draw on each member's seat allowance; Console and cloud providers bill per token. In a mixed organisation each developer is metered by whichever method they used.

SetupSee spendCap spendPer-user numbers
Claude for Teams or EnterpriseSpend report in organisation analyticsSpend limits in admin settingsSpend report CSV; Enterprise Analytics API on Enterprise
Claude Console (API)Console usage pageWorkspace spend limitsConsole Claude Code dashboard; Claude Code Analytics API
Bedrock, Agent Platform, FoundryYour cloud billing consoleYour cloud's budgetsOpenTelemetry or a gateway

OpenTelemetry export works everywhere and is the only route that streams per-user token and cost metrics into your own stack in near real time. Individual Pro and Max users have no organisation to manage; track your own credit spend, fast mode included, with /usage.

Report at contracted rates

Claude Code shows list prices by default, so if you have negotiated rates, /usage, the status line and OpenTelemetry will not match your invoice. The modelPricing managed setting fixes the reporting (it changes nothing about what Anthropic charges). Needs v2.1.242 or later.

  1. Get the rates from your contract. Per million tokens. Claude Code does not fetch them, so update the setting when the contract changes.
  2. Write the setting. Use a multiplier (below 1 for a discount, above 1 for a markup, the latter from v2.1.271), per-model overrides with the four token rates, or both. The settings reference has the exact shape.
  3. Deploy it as managed settings: server-managed, MDM, managed-settings.json or a policy helper. The key is ignored in user, project and local settings and in --settings.

Check with /usage in a session that has the managed settings: Total cost should say at your organization's configured rates. The prices in the /model picker stay at list.

Claude for Teams and Enterprise

Each member's Claude Code use draws from a per-seat allowance that resets on a rolling five-hour window and a weekly window, shared with Claude chat and Cowork. The size depends on the seat tier (Standard or Premium). Controls live in the claude.ai admin console.

  • Spend. The organisation analytics spend report shows estimated usage-credit spend per user and model, updated daily with CSV export. It appears once usage credits are on; use inside the allowance is not priced in dollars.
  • Adoption. The analytics dashboard.
  • Caps. The seat allowance is the default ceiling. Turn on usage credits to let people go further, with spend limits for the organisation, a group or an individual.
  • Per-user data. Enterprise has the Enterprise Analytics API (a Primary Owner creates a read:analytics key at claude.ai/analytics/api-keys). Teams uses the spend report CSV.

Anthropic's Claude Enterprise consumption guide gives per-user starting budgets and explains how chat, Claude Code and Cowork consume differently. Budget more for a coding seat than a chat seat: every Claude Code turn carries files, tool calls and multi-step reasoning, and one debugging session can use more than a day of chat.

Claude Console

API organisations manage spend through workspaces: set workspace spend limits on Claude Code and view cost reporting in the Console. Per member, the Console dashboard shows spend and accepted lines, and the Claude Code Analytics API returns the same daily data with an Admin API key.

Note: The first time anyone authenticates Claude Code with a Console account, a workspace called "Claude Code" is created automatically to collect all Claude Code usage. You cannot create API keys in it. If you have custom rate limits, this workspace's traffic counts towards the organisation's overall limits, so set a workspace rate limit on its Limits page to protect production workloads.

Rate limit sizing

Suggested per-user tokens per minute (TPM) and requests per minute (RPM) by organisation size:

UsersTPM per userRPM per user
1 to 5200k to 300k5 to 7
5 to 20100k to 150k2.5 to 3.5
20 to 5050k to 75k1.25 to 1.75
50 to 10025k to 35k0.62 to 0.87
100 to 50015k to 20k0.37 to 0.47
500+10k to 15k0.25 to 0.35

So 80 developers at 30k TPM each means asking for about 2.4 million TPM. The per-user figure falls as headcount rises because fewer people are active at the same moment, and limits apply to the organisation as a whole, so one person can burst above their share while others are idle. Plan extra headroom for events like live training where everyone is active at once.

Cloud providers

On Bedrock, Agent Platform and Foundry, tokens are billed to your cloud account and caps live in your cloud's billing tools. Claude Code sends no metrics back to Anthropic from these setups, so the analytics dashboards and Analytics API do not see them. For per-user attribution:

  • OpenTelemetry from each laptop to your own stack: tokens, cost and tool activity for any provider.
  • Claude apps gateway: per-user attribution, OTLP metrics and per-user spend limits.
  • An LLM gateway that tracks spend per key. Several large enterprises use the open-source LiteLLM for this; it is not affiliated with or audited by Anthropic.

Decoding a developer's limit message

MessageWhat it meansWhat to do
"You've hit your session limit" or "weekly limit"A seat usage window, shared across models; switching model will not help. The message says when it resetsRequest more with /usage-credits if credits are on, or (v2.1.234+) wait and let the task continue automatically after the reset. Control that fleet-wide with autoContinueAtUsageLimit in managed settings. See Interactive mode
"You've hit your Opus limit" or "Sonnet limit"A model-family limitSwitch to another family with /model
"individual spend limit", "org's monthly spend limit", "team's shared budget"Usage credits hit a limit you setRaise the named limit in Organisation settings, Usage, or wait for the plan reset if one is named. See Errors
A spend limit message from a Claude apps gatewayYour self-hosted gateway's capRaise the cap or wait for the period to reset
Context or auto-compact warningNot a usage limit; the conversation is near the auto-compact windowSee reducing token use
Surprisingly high API or cloud spendUsually never-cleared sessions or Opus left as defaultClear between tasks; match the model to the job

Agent teams

Agent teams run several Claude Code instances, each with its own context, so cost scales with team size and run time. With teammates in plan mode they use roughly 7x the tokens of a normal session. Use Sonnet for teammates, keep teams small and spawn prompts focused (teammates already load CLAUDE.md, MCP servers and skills), and shut teammates down when finished. Agent teams are off by default; enable them with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1.

Reduce token use

Cost tracks context size. Claude Code already helps with prompt caching for repeated content and auto-compaction near the context limit. The rest is habit.

Keep context lean

Watch usage with /usage or put it in your status line.

  • Clear between unrelated tasks. Stale context is paid for on every message. /rename first so you can /resume later.
  • Steer compaction. /compact keep the migration plan and failing test names tells Claude what to preserve. You can make that permanent in CLAUDE.md:
## When compacting
Preserve the current task list, any failing test output, and file paths we have edited. Drop exploratory reads.

Pick the right model

Sonnet handles most coding well and costs less than Opus; save Opus for architecture and long reasoning chains. Switch with /model or set a default in /config. Switching to Opus also affects subagents that inherit your model, so give simple subagents model: haiku in their definition. See Model configuration and Subagents.

Trim MCP overhead

MCP tool definitions are deferred by default, so only names and server instructions sit in context until a tool is used. Run /context to see what is taking space. CLI tools such as gh, aws, gcloud and sentry-cli are leaner still, since they add no per-tool listing. Disable servers you are not using via /mcp.

Use code intelligence for typed languages

Code intelligence plugins give Claude symbol navigation, so one "go to definition" replaces a grep plus several speculative file reads, and type errors surface automatically after edits.

Preprocess with hooks, pre-load with skills

A hook can shrink data before Claude sees it. Here is one I use on a Python project: a PreToolUse hook that rewrites pytest runs to show only the short failure summary.

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "~/.claude/hooks/quiet-pytest.sh" }]
      }
    ]
  }
}
#!/usr/bin/env bash
# ~/.claude/hooks/quiet-pytest.sh
payload=$(cat)
cmd=$(jq -r '.tool_input.command' <<<"$payload")

case "$cmd" in
  pytest*|"python -m pytest"*)
    jq --arg c "$cmd -q --tb=line --no-header 2>&1 | tail -40" \
      '{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $c})}}' <<<"$payload"
    ;;
  *)
    echo '{}'
    ;;
esac

Make it executable with chmod +x, then confirm it is registered with /hooks. To see it fire, start claude --debug-file ./debug.txt, ask for a test run, and look for a modified tool input keys line in the log.

A skill does the opposite job: it hands Claude knowledge up front (an architecture overview, naming conventions) so it does not spend tokens reading half the repo to work them out.

Move specialist instructions out of CLAUDE.md

CLAUDE.md loads into every session. Detailed procedures for PR reviews or database migrations cost tokens even when you are fixing CSS. Put them in skills, which load only when used. I keep CLAUDE.md under 200 lines.

Tune extended thinking

Thinking is on by default because it helps planning and hard reasoning, but thinking tokens bill as output and can run to tens of thousands per request. For simple work, lower the effort level with /effort or in /model, or turn thinking off in /config. Opus 5.5, Sonnet 5.5, Haiku 5.5 and the Fable models always think and cannot have it turned off. On models with a fixed thinking budget, cap it with MAX_THINKING_TOKENS (for example MAX_THINKING_TOKENS=8000); adaptive-reasoning models ignore non-zero budgets, so use effort there.

Push noisy work to subagents

Test runs, documentation fetches and log trawls produce a lot of output. A subagent keeps that in its own context and returns a summary. Its requests still cost, so give it a smaller model or run all subagents on one model.

Prompt precisely and work in small steps

"Tidy up this codebase" triggers broad scanning; "add email format validation to signup() in accounts/forms.py" does not. For larger tasks:

  • Use plan mode (Shift+Tab) so Claude proposes an approach before writing anything.
  • Press Escape as soon as it heads the wrong way, and use /rewind or double Escape to roll back (see Checkpointing).
  • Give it something to verify against: a test, a screenshot, the expected output.
  • Build and test one file at a time.

Background token use

Even idle, Claude Code spends a little: summarising past conversations for claude --resume, and status checks from commands like /usage. Typically under $0.04 a session. With prompt suggestions on, it also sends a short request after each reply to suggest your next prompt; this mostly reads from the prompt cache, is skipped near your usage limit, and can be turned off (see Interactive mode).

Why a long session gets expensive

  • Context grows. Every request carries the whole conversation, and each batch of tool results is another request. Caching makes it cheaper, not free, so a quick question in an all-day session still pays for the whole history.
  • Cache expiry. After a break longer than the cache lifetime, the next message reprocesses everything. The lifetime is an hour on a subscription, five minutes once you are drawing usage credits, and five minutes by default on API keys and cloud providers. You can choose the TTL yourself. On Pro and Max, resuming a big session after a long break offers to resume from a summary.
  • Scheduled tasks fire on schedule even when idle, each sending full context.
  • Cross-session messages arrive as new turns in an idle session. Set crossSessionInbound to hold to queue them instead (see Cross-session messaging).
  • Goal check-ins. While background work keeps a goal waiting, Claude checks in even when idle, at most three times between your prompts (v2.1.236+; uncapped before v2.1.246). Set CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 to stop them.
  • Subagents and workflows each send their own requests; workflows can spawn many.
  • Teammates keep spending until they exit.
  • Compaction reads the whole conversation it summarises, so it is a big request. If you do not need continuity, /clear is free.

The /usage breakdown flags whichever of these is taking 10% or more of recent usage.

Getting help with billing

Behaviour, including cost reporting, changes between releases; claude --version tells you what you are on. For account-specific billing questions, use the in-product messenger: on claude.ai (subscriptions) or platform.claude.com (Console), click your initials and choose Get help.