Skip to content

Prompt caching

How Claude Code caches the start of every request, which actions throw that cache away, how long it lasts, and how to check your hit rate.

Every time you send a message, Claude Code sends the model the entire conversation again: system prompt, project context, every earlier message and tool result, then your new message. The model keeps no state between requests. Prompt caching is what stops this from being painfully slow and expensive: the API recognises the part it has already processed, bills it at the cheaper cached rate, and only does full work on what is new.

Claude Code handles caching for you. You still want to know how it works, because a handful of everyday actions quietly throw the cache away and make the next turn slower and dearer. Once I learned which ones, I stopped switching models halfway through a task.

Prefix matching

The cache works on prefixes. The API compares the start of the new request with recent requests and reuses the longest exact match. On a normal turn the whole previous request matches and only the latest exchange is new.

Because matching is exact and positional, a change anywhere invalidates everything after it. There is no caching per file or per section.

Claude Code therefore orders each request from most stable to least stable:

LayerWhat it containsTypically changes when
System promptCore instructions and tool definitionsThe set of loaded tool definitions changes
Project contextCLAUDE.md, auto memory, rules without paths:A session starts, or after /clear or /compact
ConversationYour messages, Claude's replies, tool resultsEvery turn

A new turn only touches the bottom layer, so everything above stays cached. A change to the system prompt invalidates the lot.

This also explains some design choices. Plan mode and skills inject their instructions as conversation messages rather than editing the system prompt, precisely so the cached prefix survives.

Two things outside the table also matter:

  • The model. Each model has its own cache.
  • The effort level. On most models each effort level has its own cache too. The exceptions are Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1 when you use an API key or a Claude subscription.

Tip: Choose your model and effort level at the start of a session and leave them alone while you work. Save /compact for natural breaks between tasks.

Where the cache lives

The cache sits with whoever serves the model:

How you connectCache location
API key, Claude subscription, or Claude Platform on AWSAnthropic's infrastructure
Amazon Bedrock or Google CloudYour cloud provider's serving infrastructure
Microsoft FoundryAzure for "Hosted on Azure" deployments; Anthropic for "Hosted on Anthropic" deployments
Custom ANTHROPIC_BASE_URL or LLM gatewayWherever the gateway forwards requests, if it supports caching at all

Claude Code also appends system context partway through a conversation (for example, notices that a file changed on disk) and marks that block for caching on every provider, unless you set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS, in which case the block goes uncached. Bedrock (including its Mantle endpoint), Google Cloud and Foundry cache it just as the Claude API does.

Gateways and base-URL overrides such as ANTHROPIC_BEDROCK_BASE_URL behave according to how they treat the cache_control markers Claude Code sends:

  • Passed through untouched: caching behaves exactly as it would at the provider.
  • Rejected with a 400 error naming cache_control: Claude Code resends with the marker moved from the system block onto your last conversation message, and keeps doing so for the rest of the conversation. The block is billed uncached; the conversation stays cached.
  • Silently stripped: your whole history is billed as uncached input on every turn. Gateways that flatten block-style system content into a plain string cause the same problem.

See LLM gateway and data usage for more.

Actions that break the cache

Each of these causes one slow, more expensive turn while the new prefix is cached. After that, things return to normal.

Switching models

/model means the next request is read from scratch, even though the content is identical. In the terminal, Claude Code asks you to confirm a switch only while the cache is still warm (within one cache lifetime of the last request or response) and the new model is not the one that wrote the last reply. Once the cache has expired it just switches. Versions before v2.1.238 asked regardless. A PreModelSwitch hook can force or skip that confirmation.

Some switches are less obvious:

  • The opusplan setting uses Opus in plan mode and Sonnet for execution, so every plan mode toggle is a model switch.
  • Automatic model fallback on Fable models, Opus 5.5, Sonnet 5.5 and Opus 5 reruns a request flagged by a safety classifier on the fallback model, and the session carries on there.
  • A skill or command whose frontmatter sets a different model switches for that turn; your session model returns on the next prompt. A context: fork skill sets the forked subagent's model instead.

Changing effort level

On most models the next request misses the cache, and Claude Code asks you to confirm while it is warm. On Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1 with an API key or subscription, the cache survives and the change applies without a prompt. That exception does not hold on Bedrock, Google Cloud, a Claude apps gateway, with CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS set, or under a HIPAA configuration. (Before v2.1.260, Fable 5.1 lost its cache too.)

Turning on fast mode

Fast mode adds a request header that forms part of the cache key, so the first request with it misses and is billed at fast mode rates. The header is fixed when a turn starts, so turning it on mid-turn bites on the first request of your next turn. If your model does not support fast mode, enabling it also switches model, which costs a fresh cache from the next request in the current turn.

This happens once per conversation. After that, Claude Code keeps sending the header and only varies the speed setting, which is not part of the key. Turning fast mode off, the automatic drop to standard speed after a rate limit or when credits run out, and turning it on again all keep the cache. /clear and /compact reset the state because they rebuild the cache anyway.

Practical upshot: if you want fast mode, turn it on early in a session, not deep into a long one.

Connecting or removing an MCP server

Tool definitions live in the system prompt layer. What happens depends on whether tool search is deferring your MCP tools, which is the default on supported models:

  • Deferred (default): the tool list from the conversation's first request is kept for the whole conversation. Servers coming and going mid-session do not disturb the cache, and a server that finishes connecting late supplies its tools as deferred definitions.
  • Loaded up front: this applies when tool search is below its auto threshold, disabled or unavailable (for example Google Cloud models older than the Claude 4.5 generation, a custom ANTHROPIC_BASE_URL gateway, or a Foundry deployment hosted on Azure once Claude Code finds it rejects tool search).

With tools loaded up front:

Mid-session eventCacheWhat happens to definitions
A server connects, or a dynamic tool update adds toolsLostNew definitions added
A server drops on its own (a stdio process exits)KeptDefinitions stay; calls to them return errors
A remote server reconnects automaticallyKept, unless a request sent during reconnection adds the WaitForMcpServers tool for the first time, which costs one missDefinitions stay; WaitForMcpServers then stays listed
You remove a tool deliberately (a deny rule, or disabling the server in /mcp)LostDefinition removed

On resume, if a server is still connecting when the first request goes out, recorded definitions from the transcript are used so the prefix does not change when it finishes. Editing MCP config on disk changes nothing until you restart. The advisor tool is a special case: its definition sits after the cache breakpoint, so toggling /advisor is free.

Enabling or disabling a plugin

  • Skills, commands, agents, hooks, monitors and themes from a plugin never break the cache. Their content is appended after the conversation.
  • Plugins that provide MCP servers follow the MCP rules above.
  • A code intelligence plugin adds the LSP tool.

Changes made in the /plugin menu are applied through /reload-plugins when you close it, and you pay any cost on the next turn. If the reload would force a full re-read, Claude Code warns and holds off; use /reload-plugins --force to apply it anyway. Changes can also apply on their own: a plugin with a command source can reload itself, an install from /plugin may activate immediately (the summary says), /cd on v2.1.246 or later applies the new directory's plugins without the full re-read warning, and in interactive sessions (v2.1.265 or later) adding or removing a plugin in a --plugin-dir folder applies straight away, unless it would cause a full re-read, in which case you are asked to run /reload-plugins.

From v2.1.260 you can type /reload-plugins in sessions without an interactive terminal (desktop, Agent SDK, -p). There, plugin MCP server changes wait for the next session, so they never trigger a mid-session re-read.

Disabling a plugin you enabled earlier in the same session restores the previous request shape, which can hit the older cache entry if it has not expired.

Denying a whole tool

A deny rule with a bare tool name (WebFetch, Bash, the equivalent Bash(*), a name glob such as "*", or an MCP-only glob such as "mcp__*") stops Claude calling that tool from the next request, even if you add it mid-turn. With tool search active, the definitions do not change and the cache survives. Without it, the definition is removed and the cache is lost (and lost again when you remove the rule). Scoped rules such as Bash(rm *) and all allow and ask rules are checked at call time and never affect the prefix. See permissions.

Compacting

Compaction replaces history with a summary, so by design the conversation layer is rebuilt. The system prompt layer is reused (except the first compaction after resuming with a system prompt that would otherwise have changed, which switches to the current prompt once). Project context is reloaded from disk and hits the cache only if CLAUDE.md and memory have not changed since the session began.

The summary itself is produced by a request that shares your system prompt, tools and history, with a summarisation instruction on the end. While the cache is warm that request reads your prefix cheaply, so /compact costs far less than the context size suggests. After a long break, though, there is nothing to read, and the whole history is processed uncached. That is why compacting an old resumed session is the most expensive case.

Tip: If you want to abandon a line of work entirely, /rewind to an earlier turn rather than compacting. Rewinding goes back to a prefix that is already cached.

Piling up images

The API caps how many images and PDFs a request may carry, and Claude Code also caps their total size. When the next request would exceed either limit, Claude Code drops a batch of the oldest images. Claude can no longer see them; share again if needed. Removing them changes earlier messages, so the conversation is reprocessed from the earliest affected message. Because they go in batches, you get one slow turn per batch.

Upgrading Claude Code

New versions usually change the system prompt or tool definitions. Auto-update downloads in the background and applies on the next launch, never mid-session, so you see it as a slow first turn after restarting. Set DISABLE_AUTOUPDATER=1 to control timing yourself.

Actions that keep the cache

ActionWhy it is safe
Editing files in the repoFile contents only enter context when read. Claude Code appends a system reminder that a file changed and Claude re-reads if needed
Editing CLAUDE.md mid-sessionThe root and user files are read once at startup, so the edit does not apply until /clear, /compact or a restart. Nested CLAUDE.md files and paths: rules that have not loaded yet will pick up your edit when they do
Changing permission modeNo change to system prompt or tools (except opusplan, where entering plan mode switches model)
Changing output styleFrom v2.1.251 the new style arrives as a conversation message and applies from your next message. Earlier versions kept the cache but waited for /clear
Invoking skills and commandsTheir instructions are appended as user messages (unless the frontmatter sets a different model)
Running /recapAppends a summary as command output rather than replacing history
Rewinding with /rewindTruncates back to a prefix already cached; every later turn kept that entry warm. Restoring file checkpoints has no separate effect
Spawning a subagentThe call and its result are appended to your conversation

Resuming a session

Resuming resends the whole conversation, and whatever part of the prefix is unchanged and unexpired is read from cache. If your system prompt would now be different (after an upgrade, or with different --append-system-prompt text), the resumed conversation keeps its original prompt by default and adopts the new one only after compaction or in a new conversation. The CLI reference covers the cases where the prompt is rebuilt on every request instead.

How long the cache lasts

Cached prefixes expire after a period of inactivity, and each cache hit resets the clock. Stay busy and the cache stays warm; step away for long enough and your first turn back is slow. On Pro and Max, resuming a large session after a long gap prompts you to resume from a summary instead (see sessions).

There are two time-to-live options: five minutes, and one hour. The hour costs more per cache write but saves the full reprocessing when you come back from a break. For short bursts of work that never go idle, the extra write cost is wasted.

Default TTLs

Claude Code sorts each request into one of two buckets:

  • Main conversation: interactive turns, -p runs, Agent SDK turns and helpers running inline with them.
  • Everything else: subagents, workflows, in-process teammates, forks, compaction, session titles and similar side requests.
BucketClaude subscription within plan usageUsage credits, API key or cloud provider
Main conversationOne hourFive minutes
Everything elseFive minutes (a few server-controlled helpers get one hour)Five minutes

When a subscriber goes past their plan limit and starts drawing on usage credits, the main conversation drops to five minutes because the cheaper writes now matter.

Setting the TTL yourself

Each control accepts 5m or 1h; anything else is ignored. All need v2.1.242 or later.

BucketSettingEnvironment variable
Main conversationpromptCacheTtlCLAUDE_CODE_PROMPT_CACHE_TTL
Everything elsesubagentPromptCacheTtlCLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL

If you use an API key or a cloud provider and leave sessions idle often, setting promptCacheTtl to 1h is worth trying.

When several controls apply, the first match wins:

  1. FORCE_PROMPT_CACHING_5M=1 (five minutes for both buckets; handy for debugging or overriding a managed setting)
  2. The bucket's environment variable
  3. The bucket's setting
  4. For a subagent, cacheTtl inside its experimental frontmatter field (v2.1.248 or later; a 1h value is ignored while a subscription is on usage credits)
  5. ENABLE_PROMPT_CACHING_1H=1 (one hour for both buckets)
  6. The bucket default

To confirm which TTL the main conversation used, run claude -p "ping" --output-format json and look at usage.cache_creation: one-hour writes appear under ephemeral_1h_input_tokens, five-minute writes under ephemeral_5m_input_tokens.

Caveats: through a gateway set with ANTHROPIC_BASE_URL, the one-hour request relies on the anthropic-beta header, so make sure the gateway forwards it. The one-hour TTL is not available through the Claude apps gateway. On Bedrock, caching support, minimum prefix length and one-hour availability vary by model; if cache counts stay at zero, check AWS's list of supported models and regions.

Who shares a cache

In practice the cache is scoped to one machine and one directory. The system prompt embeds your auto memory paths and the conversation opens with the working directory, platform, shell and OS version, so sessions in different folders produce different prefixes. Parallel sessions in the same folder share. Sequential sessions in the same folder share only if the git status snapshot taken at startup matches.

At the API level, caches are isolated per organisation and, on some providers, per workspace. Agent SDK users running fleets can move the auto memory location out of the system prompt to share one cache across users and machines; see modifying system prompts.

Subagents, forks and the cache

A subagent has its own system prompt and tools, so it builds its own cache and does not read yours. It also sits in the "everything else" bucket, so it gets five minutes by default even on a subscription. Your cache is untouched.

A fork is different: it inherits the parent's system prompt, tools and history exactly, so its first request reads the parent's cache. The same prefix sharing applies to sessions copied with /fork (the isolation instruction is appended at the end), the compaction request, resumed subagents, and workflow fan-outs of identical agents, where Claude Code holds all but the first agent for up to five seconds by default so the rest can read the prefix the first one cached.

Checking cache performance

Every API response reports two numbers:

FieldMeaning
cache_creation_input_tokensTokens written to cache this turn, at the write rate
cache_read_input_tokensTokens served from cache, at the cheaper cached rate

Plenty of reads and little creation means caching is doing its job. Creation staying high turn after turn means something keeps changing your prefix; check the list of cache-breaking actions above.

Ways to watch:

  • Status line. A status line script can read current_usage for the per-turn numbers, and the prompt_cache object for session totals.
  • /usage. After the first response, the Session block includes a Prompt cache (main) line with hit ratio, miss count and whether the cache is warm right now (v2.1.251 or later). From v2.1.260 it also names the likely cause of the last miss, such as likely cause: tool definitions changed. See costs.
  • OpenTelemetry. For an organisation-wide view, the exporter reports cache read and creation tokens per user and session. See monitoring usage.

Turning caching off

Only worth doing when debugging a specific model or provider. Set any of these to 1:

VariableDisables caching for
DISABLE_PROMPT_CACHINGEvery model
DISABLE_PROMPT_CACHING_HAIKUThe model the haiku alias resolves to, wherever it runs (including as your main model from v2.1.283), plus a background model set through the deprecated ANTHROPIC_SMALL_FAST_MODEL
DISABLE_PROMPT_CACHING_SONNETThe model the sonnet alias resolves to
DISABLE_PROMPT_CACHING_OPUSThe model the opus alias resolves to
DISABLE_PROMPT_CACHING_FABLEFable only

The alias-based variables do not cover a different version pinned by full model ID. If sonnet points to one Sonnet release and you run another by its explicit ID, that model keeps caching; use DISABLE_PROMPT_CACHING for it. To enforce a caching policy for a whole organisation, put these or the TTL variables in the env block of managed settings.