Skip to content

Gateway compatibility guide

What Claude Code sends to an LLM gateway: API formats, headers to forward, response headers, attribution block, feature pass-through and model discovery.

This is the reference for whoever configures the gateway product. It describes what Claude Code puts on the wire, what has to reach the upstream untouched, and which features fail (loudly or silently) when it does not. Admins planning a rollout should start with Roll out an LLM gateway; developers configuring a laptop want Connect Claude Code to an LLM gateway.

Claude apps gateway is a different animal: it publishes its own endpoint reference at GET /protocol, covering its sign-in, inference, managed settings, discovery and telemetry endpoints. This page is about third-party gateways.

Two verbs are used throughout:

  • Forward unchanged: pass it upstream byte for byte.
  • Consume: the gateway may read it for routing, attribution or tracing and does not have to forward it.

Anything not marked "forward unchanged" is yours to consume or drop.

API formats

The gateway has to expose at least one of these. The client chooses one with the variables in the second column. (Google Cloud's Agent Platform is the former Vertex AI, which is why its variables still say VERTEX.)

FormatClient selects it withEndpointsMust forward unchanged
Anthropic MessagesANTHROPIC_BASE_URL/v1/messages; /v1/messages/count_tokens optionalanthropic-beta and anthropic-version headers
Amazon Bedrock InvokeModelANTHROPIC_BEDROCK_BASE_URL + CLAUDE_CODE_USE_BEDROCK=1/model/{model}/invoke, /model/{model}/invoke-with-response-stream; /model/{model}/count-tokens optionalanthropic_beta and anthropic_version body fields
Agent Platform rawPredictANTHROPIC_VERTEX_BASE_URL + CLAUDE_CODE_USE_VERTEX=1:rawPredict, :streamRawPredict; count-tokens:rawPredict optionalanthropic-beta and anthropic-version headers, plus the anthropic_version body field

Microsoft Foundry and Claude Platform on AWS both speak Anthropic Messages. Clients reach them with ANTHROPIC_FOUNDRY_BASE_URL and ANTHROPIC_AWS_BASE_URL, but a gateway in front of either implements the Anthropic Messages row. In front of Claude Platform on AWS, also forward anthropic-workspace-id, which that platform requires on every request.

Paths and incidental traffic

Match on path, not full URL. Inference posts go to /v1/messages?beta=true, and Agent Platform method suffixes hang off the publisher model path, for example /projects/{project}/locations/{location}/publishers/anthropic/models/{model}:streamRawPredict.

Token counting is the only optional piece. Without it, Claude Code estimates context use from character counts.

You will also see some best-effort traffic you can safely reject:

  • Anthropic Messages gateways get a HEAD /api/hello connection-warming probe (skipped when an HTTP proxy or client certificate is configured).
  • Bedrock-format gateways get GET /inference-profiles?type=SYSTEM_DEFINED, plus GET /inference-profiles/{profile} when the configured model is an inference profile.

Two things you will not see: the fast mode availability check and the WebFetch domain safety check both call api.anthropic.com directly, ignoring ANTHROPIC_BASE_URL. On a network that blocks that host, fast mode may report a connectivity error while inference works fine.

Gateway sessions are not eligible for the HIPAA configuration; see HIPAA setup.

Streaming rules

Claude Code consumes streaming responses event by event, so how you relay them matters:

  • Do not buffer. A gateway that waits for the full response makes Claude Code stall.
  • Deliver the whole event sequence. Claude Code expects everything through the final message_delta and message_stop. A body that ends cleanly after a content block began but before that final message_delta is treated as a dropped connection, which shows the user an incomplete-response warning or triggers a retry (see Errors).
  • Keep ping events. During long thinking pauses, pings may be the only bytes flowing. Strip or buffer them and Claude Code's streaming idle watchdog (see Network configuration) can abort mid-response. If you translate from an upstream with no pings, such as Bedrock's binary event stream, emit your own.
  • Bedrock guardrails. When a guardrail blocks a reply, Bedrock's events can reference a content block whose content_block_stop already arrived. Pass them through as sent.
  • Bedrock format stays binary. On the InvokeModel format, Claude Code expects invoke-with-response-stream as application/vnd.amazon.eventstream. Converting it to SSE or changing the Content-Type makes it unparseable; Amazon Bedrock describes the symptom.

Format mismatches

The commonest bug is a client speaking one format to the gateway while the upstream accepts another. With the Bedrock or Agent Platform format, Claude Code sends only the subset those providers accept. With the Anthropic Messages format it sends everything, even if your gateway then forwards to Bedrock. Bridging that gap is the gateway's job. If your upstream is Bedrock or Agent Platform, the simplest fix is often to expose that provider's own format and configure clients as the connect page shows.

How the connection method changes what the client sends

There are three client behaviours a gateway can see:

  1. Provider format. CLAUDE_CODE_USE_BEDROCK=1 with ANTHROPIC_BEDROCK_BASE_URL, or CLAUDE_CODE_USE_VERTEX=1 with ANTHROPIC_VERTEX_BASE_URL. Claude Code uses that provider's IDs, fields and defaults.
  2. Anthropic Messages. ANTHROPIC_BASE_URL pointing at you. Claude Code treats you as the Claude API and has no idea what is upstream.
  3. Claude apps gateway sign-in. Anthropic Messages on the wire, but because that gateway can route anywhere, Claude Code restricts itself to the beta values and capability assumptions Bedrock and Agent Platform also accept.
BehaviourProvider formatAnthropic MessagesClaude apps gateway
Default model IDsProvider form, e.g. us.anthropic.claude-opus-4-8 on BedrockAnthropic form, e.g. claude-opus-4-8Anthropic form
anthropic-beta valuesSubset Bedrock and Agent Platform acceptFull set, unless CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS is setBedrock and Agent Platform subset
Fields for an unrecognised ID (such as an alias)Fixed thinking budget; no effort or context managementEverything current models accept on the Claude API: adaptive reasoning, effort, context managementSame as provider format
One-hour prompt cache TTL when opted inttl in cache_control, no beta valuettl plus extended-cache-ttl in anthropic-beta, which you must forwardSee Claude apps gateway
Background task model unless ANTHROPIC_DEFAULT_HAIKU_MODEL is setDefault Sonnet, or the main model once one is chosenMain model, or default Haiku when an Anthropic Console key is supplied via ANTHROPIC_API_KEY or apiKeyHelper and ANTHROPIC_AUTH_TOKEN is unsetMain model

Foundry and Claude Platform on AWS are left out of that table; see Microsoft Foundry and Claude Platform on AWS. Feature availability and Data usage cover what each connection supports and what telemetry it sends by default.

Unrecognised model IDs

Whatever the connection method, two client-side settings fix assumptions about IDs Claude Code does not know:

  • Context window. Assumed to be 200K, or 1M if the ID contains [1m]. Declare the real figure as described in Model configuration.
  • Capabilities. Map the underlying model's Anthropic ID to your alias with a modelOverrides entry in the settings you distribute.

Request headers

Header names are case-insensitive. Forward anthropic-version and anthropic-beta unchanged, and anthropic-workspace-id for Claude Platform on AWS. Everything else is yours to consume.

HeaderWhat it carries
Authorization / x-api-keyThe developer's gateway credential, in one or both depending on the client's credential variable
anthropic-versionCurrently 2023-06-01. Provider-format requests also carry an anthropic_version body field holding the provider dialect string, which is a different value
anthropic-betaComma-separated capability flags. Forward verbatim; do not allowlist values. When the developer is on a claude.ai login (base URL set, no gateway credential), it also carries an OAuth capability the upstream needs, and stripping it gives 401
x-claude-code-session-idUnique per session. The cheapest way to group requests without parsing bodies
x-claude-code-agent-idThe subagent that sent the request; only present for spawned agents
x-claude-code-parent-agent-idThe agent that spawned the requester; nested agents only

Subagent IDs are new each spawn; agent team teammates reuse a stable name-based ID. Either way it identifies an agent, never a person or device. ANTHROPIC_CUSTOM_HEADERS set by developers also appear.

Gateway hint headers

From v2.1.273, Claude Code can attach routing hints describing each request. Defaults:

  • Direct to the Anthropic API: sent.
  • Custom base URL: off, because a strict proxy might reject unknown headers. Turn on with CLAUDE_CODE_GATEWAY_HINT_HEADERS=1, ideally in managed settings.
  • Bedrock, Agent Platform, Foundry, Claude Platform on AWS: only when that variable is 1.

Setting it to 0 stops them everywhere. Values are printable ASCII and never contain prompt text or file contents.

HeaderValues and meaning
x-claude-code-request-classmain, subagent, workflow, compaction or auxiliary (titles, classifiers, summaries). On every request
x-claude-code-agent-typeBuilt-in type such as Explore, Plan or general-purpose; or custom, teammate, fork. Only on a subagent's own turns, never the user's chosen agent name
x-claude-code-compactionOn the summarising request only: auto, manual (from /compact) or reactive (after a too-long rejection)
x-claude-code-context-compactedOnce, on the first main request after a compaction, same values. The earlier prefix is dead, so caches keyed on it can be evicted
x-claude-code-prev-tool-durationsRun times of the tool calls whose results this request carries, for example Grep=31;Bash=2210
x-claude-code-prompt-idRandom UUID shared by every request serving one user prompt, subagents included (v2.1.283+)

If you parse the tool durations header: entries are whole milliseconds in result-collection order; at most 32 entries and 4 KB, keeping the earliest; tool names are percent-encoded (%, ;, =, comma, space and non-printable ASCII); split on ; then = and decode. It never appears on compaction calls, side requests or the first request of a prompt, so absence does not mean no tools ran. Times exclude permission prompts and hooks, and parallel calls each report their own time, so they will not sum to the gap between requests.

Treat everything as an open list

Claude Code grows new anthropic-beta values, body fields and occasionally new anthropic-* or x-claude-code-* headers every few releases. Towards an Anthropic-format upstream, pass anthropic-* headers and body fields through wholesale. An allowlist built from today's traffic will strip tomorrow's feature on the day it ships. The only exception is a non-Anthropic upstream such as Bedrock or Agent Platform, where translating between schemas is the gateway's responsibility.

Response headers

HeaderWhat to return
content-typetext/event-stream on streamed Anthropic responses; application/vnd.amazon.eventstream, unmodified, on Bedrock-format responses
retry-afterInteger seconds, not an HTTP date. Claude Code waits at least that long before retrying, and outside CLAUDE_CODE_RETRY_WATCHDOG sessions a value over 60 stops retries and shows the error immediately
x-should-retryPass the upstream value through. true marks retryable, false not
anthropic-ratelimit-unified-*Pass through on every response. Used to show plan usage to claude.ai users and to tell a plan limit or spend cap from a temporary throttle on 429

Also forward error bodies unmodified; see error forwarding.

The attribution block in the system prompt

Claude Code prepends a short block to the system prompt holding the client version and a fingerprint derived from the conversation. api.anthropic.com strips it before processing, provided it arrives unchanged as the first system block, so first-party prompt caching is unaffected. Any other upstream sees it as prompt text.

Because the strip is positional:

  • Forward the system array exactly as received, block first. Prepending a block, reordering, or flattening to a string defeats the strip, and the block then reaches the model and the cache key.
  • Keep it in its own array entry. A merged block starting with the attribution header is treated as attribution in full and everything merged into it, the rest of the system prompt included, is dropped.
  • If you must reshape system content, set CLAUDE_CODE_ATTRIBUTION_HEADER=0 on clients so the block is never sent. Do not strip or move it in the gateway, since Anthropic and the cloud providers read it for attribution.

That variable is a caching-compatibility switch, not a privacy control. One exception: when requests go to api.anthropic.com (no third-party provider, base URL unset or naming that host) and the credential is not an Anthropic profile or federation credential, the block stays on auto mode classifier requests even with the variable at 0, because it is the only marker identifying them as Claude Code traffic. Through a gateway, on a cloud provider, or with a profile or federation credential, 0 removes it there too. (Before v2.1.229 there was no exception and auto mode failed when the API declined the unidentified requests.)

Since v2.1.181 the block is stable for a conversation's lifetime behind a custom base URL, so body-keyed caches and upstream prompt caching work. On older versions it contained a per-request token; set CLAUDE_CODE_ATTRIBUTION_HEADER=0 there if your gateway caches on the request body or forwards to Bedrock, Foundry or Agent Platform.

Feature pass-through

Behind ANTHROPIC_BASE_URL, Claude Code sends essentially what it would send to api.anthropic.com, minus a small, release-dependent set of direct-only diagnostics and defaults. Fine-grained tool streaming is one such default: off behind a custom base URL unless developers set CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING=1.

The important rule: features that add body fields pair them with a beta header, and the pair must travel together. Strip the header but keep the body, or forward an Anthropic body to an upstream with a different schema, and you get hard 400s. Only when both halves are missing does the feature quietly switch off. Redacting or rewriting bodies for inspection breaks pairs in the same way, so inspect without modifying.

FeatureWhat travelsFailure symptomFix
Adaptive reasoningNo beta header; thinking: {"type": "adaptive"} for Claude 4.6+, and for unrecognised names such as aliases400 naming thinking or adaptiveUpgrade the upstream; on Opus 4.6 and Sonnet 4.6, CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1
Context managementBeta header plus context_management field400 with Extra inputs are not permitted, typically Anthropic in, Bedrock outForward both, or disable pre-release capabilities
Extended context, interleaved thinkingHeader onlySilently unavailableForward anthropic-beta verbatim
Beta tool fields (strict, defer_loading)Header plus tool schema fields400 naming the unknown fieldForward both, or disable pre-release capabilities
Effort, structured outputs, task budgetoutput_config plus a header each400 naming output_config on Bedrock or Agent PlatformForward together; the experimental-betas switch drops format and task budget but not effort; CLAUDE_CODE_DISABLE_STRUCTURED_OUTPUTS=1 (v2.1.288+) drops only the format
Prompt cachingcache_control on system blocks and messages entries, including mid-conversation role: "system" entriesNo error; every turn bills as uncached inputForward cache_control wherever it appears; do not flatten block content to strings
Token countingcount_tokens endpointNo error; /context shows estimatesExpose the endpoint

The ANTHROPIC_DEFAULT_*_MODEL_SUPPORTED_CAPABILITIES variables only work with CLAUDE_CODE_USE_BEDROCK, CLAUDE_CODE_USE_VERTEX, CLAUDE_CODE_USE_FOUNDRY or CLAUDE_CODE_USE_MANTLE. They do nothing behind an ANTHROPIC_BASE_URL gateway.

Retries and error forwarding

How Claude Code reacts to an upstream rejection:

  • thinking field, a mid-conversation system message, or its cache_control rejected: retries and disables that capability for the rest of the conversation.
  • Thinking signature rejected (including a 400 saying the block is bound to a different conversation): drops earlier thinking blocks, retries, and keeps them out from then on. New replies still think. That particular error comes from the preserved-thinking check, which fails if system, tools or earlier messages differ from the original request, so a gateway that rewrites any of those can cause it.
  • Advisor tool entry rejected as an unknown tool type (a 400 or 422 whose message has the type after Input tag, such as Input tag 'advisor_20260301'): retries once without it and its beta value, leaves the advisor out for that base URL until exit, and /advisor is unavailable meanwhile. From v2.1.280.
  • output_config.effort rejected (a 400 naming it with Extra inputs are not permitted, or saying the model does not support effort): retries without effort and omits it for that model until exit.
  • Context management or tool schema fields rejected: no retry; the developer sees the 400.

All of this matches on the upstream's wording, so return error bodies unmodified. Wrapping errors in your own envelope breaks recovery even if the status code survives, unless the envelope's message carries a stable capability_rejected: token such as capability_rejected: prompt_too_long, the scheme Claude apps gateway uses.

Disabling pre-release capabilities

If you cannot forward both halves of the beta pairs, have developers set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1.

It removes:

  • context management and context_management;
  • beta tool fields such as strict and defer_loading (standard name, description, input_schema and cache_control stay);
  • output_config.format (v2.1.287+) and output_config.task_budget;
  • MCP tool search, so all MCP tools load upfront, unless managed settings keep it on.

It keeps:

  • the extended context, interleaved thinking and effort beta values, which cloud providers accept;
  • output_config.effort and the adaptive thinking field;
  • the OAuth beta value subscription auth needs;
  • anything developers add through ANTHROPIC_BETAS or CLAUDE_CODE_EXTRA_BODY.

If a host platform embedding Claude Code sets CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST, the switch does not stop auto mode sessions on Bedrock, Agent Platform, Foundry or a Claude apps gateway from requesting server-side classifier review, which adds a beta value and a safeguards field. CLAUDE_CODE_AUTO_MODE_SERVER=0 stops that.

From v2.1.227 an organisation can keep MCP tool search on under the switch through managed settings. Direct or via ANTHROPIC_BASE_URL, Claude Code then keeps the tool-search beta header, defer_loading fields and tool_reference blocks and strips the rest. On a cloud provider or a Claude apps gateway sign-in, the override does nothing.

Model discovery

With an Anthropic Messages gateway, Claude Code can call the gateway's /v1/models at startup and add the results to /model. It is opt-in via CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1, because a gateway backed by a shared key would otherwise show everyone every model that key can reach. A modelPicker lineup with replaceBuiltInOptions hides discovered models.

When it runs

Never when any CLAUDE_CODE_USE_* variable is set (even alongside ANTHROPIC_BASE_URL), or when the base URL is unset or api.anthropic.com. It does run with nonessential traffic disabled, since it only talks to your gateway (from v2.1.257).

The request

GET /v1/models?limit=1000 with a 3-second default timeout, raised with CLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MS (v2.1.269+). Any redirect, even http to https, is a silent failure, so serve it at the configured base URL.

Credentials (both headers from v2.1.248; earlier versions sent one):

  • Authorization: ANTHROPIC_AUTH_TOKEN as bearer, otherwise the apiKeyHelper value.
  • x-api-key: the resolved API key such as ANTHROPIC_API_KEY; if a helper is the only credential, its value goes here too.

ANTHROPIC_CUSTOM_HEADERS are included, and a non-empty custom header replaces a built-in header of the same name (case-insensitive). If neither credential header resolves, discovery is skipped with a [gatewayDiscovery] skipped debug line, even if a credential is supplied only through custom headers.

The response

Claude Code reads id, plus optional display_name and description, from each item in data:

{
  "data": [
    { "id": "claude-opus-4-8", "display_name": "Opus (EU region)", "description": "Use for design reviews and gnarly refactors" },
    { "id": "bedrock/anthropic.claude-haiku-4-5" },
    { "id": "internal-summariser-v2" }
  ]
}

An entry is kept if its id contains claude or anthropic anywhere, case-insensitive. In the example, the first two are kept and the third ignored. (Before v2.1.223 the ID had to start with one of those words.)

How entries appear

  • Name: display_name if it differs from the id; otherwise the model's name if Claude Code recognises the ID; otherwise the raw id. So acme-claude-sonnet-4-6 with no display name shows as Sonnet 4.6.
  • Description: collapsed to one line; "From gateway" if absent.
  • Filtering: only models permitted by the availableModels managed setting are added.
  • De-duplication: no separate row if the ID matches an existing row (or is another spelling of the same Fable version), or if it is the exact model a built-in alias currently resolves to. While sonnet resolves to claude-sonnet-5-5, a discovered claude-sonnet-5-5 folds into the sonnet row, while claude-sonnet-5 gets its own.

Results are cached in ~/.claude/cache/gateway-models.json (%USERPROFILE%\.claude\cache\gateway-models.json on Windows, or under CLAUDE_CONFIG_DIR if set) and refreshed each startup. On failure, the picker uses the previous cache or the built-in list. Aliases that fail the name filter can be added by hand with the model configuration variables.