Manage costs
Track Claude Code spend with /usage and /insights, control it across Teams, Enterprise, Console and cloud setups, and cut token use day to day.
Claude Code is metered in tokens. Subscription plans (Pro, Max, Team, Enterprise) wrap that in a plan allowance; API and cloud-provider access bills each token. Either way, what you spend depends mostly on the model you pick, how big your context gets, and whether you run several sessions or automation in parallel.
For a planning figure: across enterprise deployments the average is about $13 per developer per active day, or $150 to $250 per developer per month, and 90% of users stay under $30 per active day. I always tell clients to run a small pilot first and use the tools below to get their own baseline before committing to a number.
This page covers Claude Code only. Usage limits for other Claude products are documented in Anthropic's Help Centre.
Track your own usage
/usage
The Session block at the top of /usage breaks down tokens for the current session:
Total cost: $1.92
Total duration (API): 14m 05s
Total duration (wall): 2h 41m 37s
Total code changes: 312 lines added, 88 lines removed
Usage by model:
claude-opus-4-8: 3.4k input, 18.9k output, 2.1m cache read, 140.0k cache write ($1.81)
claude-haiku-4-5: 9.2k input, 1.1k output ($0.11)
Things to know about that figure:
- It is computed on your machine from token counts at list price, so it is an estimate. The Usage page in the Claude Console is the authoritative bill.
- If an admin has set a
modelPricingtable in managed settings, the figure uses your contracted rates and theTotal costline saysat your organization's configured rates(see below). - Responses billed at the 1.1x data residency rate are multiplied by 1.1. The same figure feeds the status line cost field and counts towards
--max-budget-usd(see CLI reference). /clearstarts a new session, so totals reset to zero. Before v2.1.211 they accumulated for the life of the process.- For Pro and Max subscribers the dollar figure is not what you pay; your usage is in the plan. You also see plan usage bars, activity stats and a breakdown on the same screen.
Prompt cache line
From v2.1.251, after the first response a Prompt cache (main) line summarises how well prompt caching is working: request count, share of input tokens served from cache, misses, and whether the cache is currently warm. For example:
Prompt cache (main): 22 requests · 88% of input tokens from cache · 3 misses (last 12m ago, 205.4k tokens re-cached, likely cause: tool definitions changed) · warm (5m TTL, last activity 1m ago)
- Misses are requests that reprocessed content the cache already held, with when the last one happened and how many tokens were re-written. From v2.1.260 a likely cause is named when Claude Code can tell.
- Expected rebuilds appear once compaction or tool-result clearing has rewritten the conversation; those misses are counted separately because Claude Code caused them on purpose.
- Warm or cold says whether the cached prefix is still within its lifetime and which TTL applies. Cold shows idle time. If no response reported cache tokens, the line ends
no prompt caching reported by the API.
It reads the cache fields in API responses, so it works on any provider or gateway. It covers the main conversation, not subagents, and resets with /clear. Status line scripts get the same numbers from the prompt_cache object.
Plan usage breakdown
On Pro, Max, Team and Enterprise, /usage also explains what is eating your allowance:
- Attribution: percentage of recent usage from skills, subagents, plugins and individual MCP servers. An MCP server is only charged for requests that consumed one of its tool results (before v2.1.222 it was over-counted).
- Behaviour flags: anything such as long context or cache misses that accounts for 10% or more of recent usage, each with a tip.
- Loops (v2.1.242+): the heaviest
/loopand other scheduled tasks, with how often each fires, run count, total and per-run tokens, and last run. Rows are keyed on the prompt, so stopping and recreating a loop keeps one row.
Press d or w to flip between 24 hours and 7 days. Figures are approximate and come from local history, so other devices and claude.ai are excluded. In VS Code, the attribution and flags appear in the Account & usage dialog with a Day/Week toggle, without Loops.
Usage-credits row
While usage credits are on, /usage shows a row for them:
- Pro and Max: this month's spend against your monthly limit, or
Unlimitedwith no figure if you have not set one. - Team and Enterprise: your own spend this month against any limit that applies to you personally. Organisation-wide limits are not shown here. No personal limit means spend with nothing beside it. If credits are off for you, there is no row.
With a limit set, the row appears straight away at 0% (before v2.1.236 it only appeared on Pro and Max, and only after first spend).
If the usage request fails
Usually because the usage endpoint is rate limited. /usage then shows the last bars it loaded on this machine in the past 60 minutes with a Showing last-known usage note and the data's age. Press r to retry. With no recent snapshot it reports the rate limit and offers the same retry.
/insights
/insights is about how you work, not how much. It analyses recent sessions on this machine and writes an HTML report on what you work on, where things went wrong (misunderstood requests, buggy code) and how to get more out of Claude Code. Each run takes up to 200 sessions it has not seen before and skips very short ones; if some are left out, the header shows something like 200 sessions (412 total). If auto mode is available and you mostly worked without it, the report can estimate how many permission prompts it would have handled.
The latest report is ~/.claude/usage-data/report.html, with a timestamped copy per run alongside. Reports are cleaned up with other session data after cleanupPeriodDays (30 days by default). It works on any plan and provider, uses your normal account so it costs tokens, and ignores other devices and claude.ai. See Commands.
Usage credits on a subscription
Usage credits let you keep going past your plan limit. Run /usage-credits while signed in with a claude.ai subscription (it does not work with API key auth). Self-serve Enterprise, Enterprise trials and Enterprise billed through AWS Marketplace need v2.1.248 or later; older versions report Unknown command: /usage-credits.
| You are | /usage-credits does |
|---|---|
| A Pro or Max subscriber | Opens Settings, Usage on claude.ai, where you switch credits on or off and see balance, monthly spend and limit |
| Team or Enterprise with billing access | Opens Organisation settings, Usage |
| Team or Enterprise without billing access | Asks you to confirm, then sends a request to your admins |
The request flow only works interactively; with -p or from Remote Control it tells you to use an interactive session. Repeating it while a request is pending tells you one is already waiting; once an admin dismisses it, you can send another (before v2.1.222 a dismissal blocked new requests). On Pro and Max, hitting your spend limit with credits left prompts you to raise or remove the limit without leaving the CLI.
Manage costs for an organisation
Your controls depend on how people sign in. Teams and Enterprise draw on each member's seat allowance; Console and cloud providers bill per token. In a mixed organisation each developer is metered by whichever method they used.
| Setup | See spend | Cap spend | Per-user numbers |
|---|---|---|---|
| Claude for Teams or Enterprise | Spend report in organisation analytics | Spend limits in admin settings | Spend report CSV; Enterprise Analytics API on Enterprise |
| Claude Console (API) | Console usage page | Workspace spend limits | Console Claude Code dashboard; Claude Code Analytics API |
| Bedrock, Agent Platform, Foundry | Your cloud billing console | Your cloud's budgets | OpenTelemetry or a gateway |
OpenTelemetry export works everywhere and is the only route that streams per-user token and cost metrics into your own stack in near real time. Individual Pro and Max users have no organisation to manage; track your own credit spend, fast mode included, with /usage.
Report at contracted rates
Claude Code shows list prices by default, so if you have negotiated rates, /usage, the status line and OpenTelemetry will not match your invoice. The modelPricing managed setting fixes the reporting (it changes nothing about what Anthropic charges). Needs v2.1.242 or later.
- Get the rates from your contract. Per million tokens. Claude Code does not fetch them, so update the setting when the contract changes.
- Write the setting. Use a
multiplier(below 1 for a discount, above 1 for a markup, the latter from v2.1.271), per-modeloverrideswith the four token rates, or both. The settings reference has the exact shape. - Deploy it as managed settings: server-managed, MDM,
managed-settings.jsonor a policy helper. The key is ignored in user, project and local settings and in--settings.
Check with /usage in a session that has the managed settings: Total cost should say at your organization's configured rates. The prices in the /model picker stay at list.
Claude for Teams and Enterprise
Each member's Claude Code use draws from a per-seat allowance that resets on a rolling five-hour window and a weekly window, shared with Claude chat and Cowork. The size depends on the seat tier (Standard or Premium). Controls live in the claude.ai admin console.
- Spend. The organisation analytics spend report shows estimated usage-credit spend per user and model, updated daily with CSV export. It appears once usage credits are on; use inside the allowance is not priced in dollars.
- Adoption. The analytics dashboard.
- Caps. The seat allowance is the default ceiling. Turn on usage credits to let people go further, with spend limits for the organisation, a group or an individual.
- Per-user data. Enterprise has the Enterprise Analytics API (a Primary Owner creates a
read:analyticskey atclaude.ai/analytics/api-keys). Teams uses the spend report CSV.
Anthropic's Claude Enterprise consumption guide gives per-user starting budgets and explains how chat, Claude Code and Cowork consume differently. Budget more for a coding seat than a chat seat: every Claude Code turn carries files, tool calls and multi-step reasoning, and one debugging session can use more than a day of chat.
Claude Console
API organisations manage spend through workspaces: set workspace spend limits on Claude Code and view cost reporting in the Console. Per member, the Console dashboard shows spend and accepted lines, and the Claude Code Analytics API returns the same daily data with an Admin API key.
Note: The first time anyone authenticates Claude Code with a Console account, a workspace called "Claude Code" is created automatically to collect all Claude Code usage. You cannot create API keys in it. If you have custom rate limits, this workspace's traffic counts towards the organisation's overall limits, so set a workspace rate limit on its Limits page to protect production workloads.
Rate limit sizing
Suggested per-user tokens per minute (TPM) and requests per minute (RPM) by organisation size:
| Users | TPM per user | RPM per user |
|---|---|---|
| 1 to 5 | 200k to 300k | 5 to 7 |
| 5 to 20 | 100k to 150k | 2.5 to 3.5 |
| 20 to 50 | 50k to 75k | 1.25 to 1.75 |
| 50 to 100 | 25k to 35k | 0.62 to 0.87 |
| 100 to 500 | 15k to 20k | 0.37 to 0.47 |
| 500+ | 10k to 15k | 0.25 to 0.35 |
So 80 developers at 30k TPM each means asking for about 2.4 million TPM. The per-user figure falls as headcount rises because fewer people are active at the same moment, and limits apply to the organisation as a whole, so one person can burst above their share while others are idle. Plan extra headroom for events like live training where everyone is active at once.
Cloud providers
On Bedrock, Agent Platform and Foundry, tokens are billed to your cloud account and caps live in your cloud's billing tools. Claude Code sends no metrics back to Anthropic from these setups, so the analytics dashboards and Analytics API do not see them. For per-user attribution:
- OpenTelemetry from each laptop to your own stack: tokens, cost and tool activity for any provider.
- Claude apps gateway: per-user attribution, OTLP metrics and per-user spend limits.
- An LLM gateway that tracks spend per key. Several large enterprises use the open-source LiteLLM for this; it is not affiliated with or audited by Anthropic.
Decoding a developer's limit message
| Message | What it means | What to do |
|---|---|---|
| "You've hit your session limit" or "weekly limit" | A seat usage window, shared across models; switching model will not help. The message says when it resets | Request more with /usage-credits if credits are on, or (v2.1.234+) wait and let the task continue automatically after the reset. Control that fleet-wide with autoContinueAtUsageLimit in managed settings. See Interactive mode |
| "You've hit your Opus limit" or "Sonnet limit" | A model-family limit | Switch to another family with /model |
| "individual spend limit", "org's monthly spend limit", "team's shared budget" | Usage credits hit a limit you set | Raise the named limit in Organisation settings, Usage, or wait for the plan reset if one is named. See Errors |
| A spend limit message from a Claude apps gateway | Your self-hosted gateway's cap | Raise the cap or wait for the period to reset |
| Context or auto-compact warning | Not a usage limit; the conversation is near the auto-compact window | See reducing token use |
| Surprisingly high API or cloud spend | Usually never-cleared sessions or Opus left as default | Clear between tasks; match the model to the job |
Agent teams
Agent teams run several Claude Code instances, each with its own context, so cost scales with team size and run time. With teammates in plan mode they use roughly 7x the tokens of a normal session. Use Sonnet for teammates, keep teams small and spawn prompts focused (teammates already load CLAUDE.md, MCP servers and skills), and shut teammates down when finished. Agent teams are off by default; enable them with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1.
Reduce token use
Cost tracks context size. Claude Code already helps with prompt caching for repeated content and auto-compaction near the context limit. The rest is habit.
Keep context lean
Watch usage with /usage or put it in your status line.
- Clear between unrelated tasks. Stale context is paid for on every message.
/renamefirst so you can/resumelater. - Steer compaction.
/compact keep the migration plan and failing test namestells Claude what to preserve. You can make that permanent in CLAUDE.md:
## When compacting
Preserve the current task list, any failing test output, and file paths we have edited. Drop exploratory reads.
Pick the right model
Sonnet handles most coding well and costs less than Opus; save Opus for architecture and long reasoning chains. Switch with /model or set a default in /config. Switching to Opus also affects subagents that inherit your model, so give simple subagents model: haiku in their definition. See Model configuration and Subagents.
Trim MCP overhead
MCP tool definitions are deferred by default, so only names and server instructions sit in context until a tool is used. Run /context to see what is taking space. CLI tools such as gh, aws, gcloud and sentry-cli are leaner still, since they add no per-tool listing. Disable servers you are not using via /mcp.
Use code intelligence for typed languages
Code intelligence plugins give Claude symbol navigation, so one "go to definition" replaces a grep plus several speculative file reads, and type errors surface automatically after edits.
Preprocess with hooks, pre-load with skills
A hook can shrink data before Claude sees it. Here is one I use on a Python project: a PreToolUse hook that rewrites pytest runs to show only the short failure summary.
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [{ "type": "command", "command": "~/.claude/hooks/quiet-pytest.sh" }]
}
]
}
}
#!/usr/bin/env bash
# ~/.claude/hooks/quiet-pytest.sh
payload=$(cat)
cmd=$(jq -r '.tool_input.command' <<<"$payload")
case "$cmd" in
pytest*|"python -m pytest"*)
jq --arg c "$cmd -q --tb=line --no-header 2>&1 | tail -40" \
'{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $c})}}' <<<"$payload"
;;
*)
echo '{}'
;;
esac
Make it executable with chmod +x, then confirm it is registered with /hooks. To see it fire, start claude --debug-file ./debug.txt, ask for a test run, and look for a modified tool input keys line in the log.
A skill does the opposite job: it hands Claude knowledge up front (an architecture overview, naming conventions) so it does not spend tokens reading half the repo to work them out.
Move specialist instructions out of CLAUDE.md
CLAUDE.md loads into every session. Detailed procedures for PR reviews or database migrations cost tokens even when you are fixing CSS. Put them in skills, which load only when used. I keep CLAUDE.md under 200 lines.
Tune extended thinking
Thinking is on by default because it helps planning and hard reasoning, but thinking tokens bill as output and can run to tens of thousands per request. For simple work, lower the effort level with /effort or in /model, or turn thinking off in /config. Opus 5.5, Sonnet 5.5, Haiku 5.5 and the Fable models always think and cannot have it turned off. On models with a fixed thinking budget, cap it with MAX_THINKING_TOKENS (for example MAX_THINKING_TOKENS=8000); adaptive-reasoning models ignore non-zero budgets, so use effort there.
Push noisy work to subagents
Test runs, documentation fetches and log trawls produce a lot of output. A subagent keeps that in its own context and returns a summary. Its requests still cost, so give it a smaller model or run all subagents on one model.
Prompt precisely and work in small steps
"Tidy up this codebase" triggers broad scanning; "add email format validation to signup() in accounts/forms.py" does not. For larger tasks:
- Use plan mode (Shift+Tab) so Claude proposes an approach before writing anything.
- Press Escape as soon as it heads the wrong way, and use
/rewindor double Escape to roll back (see Checkpointing). - Give it something to verify against: a test, a screenshot, the expected output.
- Build and test one file at a time.
Background token use
Even idle, Claude Code spends a little: summarising past conversations for claude --resume, and status checks from commands like /usage. Typically under $0.04 a session. With prompt suggestions on, it also sends a short request after each reply to suggest your next prompt; this mostly reads from the prompt cache, is skipped near your usage limit, and can be turned off (see Interactive mode).
Why a long session gets expensive
- Context grows. Every request carries the whole conversation, and each batch of tool results is another request. Caching makes it cheaper, not free, so a quick question in an all-day session still pays for the whole history.
- Cache expiry. After a break longer than the cache lifetime, the next message reprocesses everything. The lifetime is an hour on a subscription, five minutes once you are drawing usage credits, and five minutes by default on API keys and cloud providers. You can choose the TTL yourself. On Pro and Max, resuming a big session after a long break offers to resume from a summary.
- Scheduled tasks fire on schedule even when idle, each sending full context.
- Cross-session messages arrive as new turns in an idle session. Set
crossSessionInboundtoholdto queue them instead (see Cross-session messaging). - Goal check-ins. While background work keeps a goal waiting, Claude checks in even when idle, at most three times between your prompts (v2.1.236+; uncapped before v2.1.246). Set
CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0to stop them. - Subagents and workflows each send their own requests; workflows can spawn many.
- Teammates keep spending until they exit.
- Compaction reads the whole conversation it summarises, so it is a big request. If you do not need continuity,
/clearis free.
The /usage breakdown flags whichever of these is taking 10% or more of recent usage.
Getting help with billing
Behaviour, including cost reporting, changes between releases; claude --version tells you what you are on. For account-specific billing questions, use the in-product messenger: on claude.ai (subscriptions) or platform.claude.com (Console), click your initials and choose Get help.