Skip to content

Claude apps gateway configuration

Every gateway.yaml option for the Claude apps gateway, from listener, OIDC and Postgres through upstreams, model routing, group policies and telemetry.

A Claude apps gateway is configured entirely by one YAML file, usually called gateway.yaml. It says where the gateway listens, how people sign in, where inference goes, and which policies and telemetry apply to whom. This page is the reference for that file, section by section.

If you have not run a gateway yet, start with the walkthrough in Claude apps gateway, then come back here as the config grows. Hosting and operations are in the deployment guide.

The file is read once, at startup:

claude gateway --config /etc/claude/gateway.yaml

Every field is validated against a schema at boot, and unknown keys are rejected. A typo stops the server with a named error rather than being silently ignored, which I have come to appreciate after one too many YAML indentation mistakes.

Shape of the file

SectionRequiredPurpose
listenYesBind address, public origin, TLS
oidcYesIdentity provider, claims and who may sign in
sessionYesSigning secret and lifetime of the gateway's bearer tokens
storeYesPostgreSQL connection
upstreamsYesModel providers, in failover order
adminNoAdmin API access and retention for spend limits
enforcementNoFail-open or fail-closed spend checks
pricingNoContracted rates and a multiplier for the spend meter and client cost displays
models, auto_include_builtin_modelsNoCurated model list and per-upstream model IDs
managedNoPolicies (managed settings and model access) by IdP group or email domain
telemetryNoOTLP forwarding to your collectors
access_control, limits, timeouts, rate_limitsNoIP filtering, size caps, upstream wait, sign-in rate limits
load_test_modeNoCanned replies for load testing without a provider

Keeping secrets out of the file

Never paste a client secret, signing key or database password into gateway.yaml. Reference it instead, and the gateway resolves the value at boot:

SyntaxResolves toGood for
${NAME}Environment variable NAME. Boot fails if it is unset.Container env vars, secrets injected by your orchestrator
${file:/abs/path}Trimmed contents of that fileKubernetes Secret mounts, Vault Agent, SOPS

${NAME} can be embedded inside a longer string. ${file:...} must be the entire value of the field, so you cannot splice it into a URL. For a database password stored in a file, use store.password rather than putting it inside postgres_url.

Required sections

listen

KeyDefaultNotes
host0.0.0.0Bind address.
port8080Bind port.
public_urlnoneThe external https:// origin. Required unless host is loopback. Used to build the IdP redirect URI, the discovery document and the OTLP endpoint pushed to clients, so telemetry also needs it. The gateway never derives its origin from X-Forwarded-* headers because clients can forge them.
tls.cert, tls.keynonePEM file paths, if the gateway terminates TLS itself.
trusted_proxiesnoneCIDRs or IPs of load balancers in front. X-Forwarded-For is honoured only from these, so per-IP rate limits, allow lists and audit see real client addresses. Entries formatted ipv4:port or [ipv6]:port are read with the port dropped; an unbracketed IPv6 address with a port may be misread, so disable that format on the proxy.

oidc

This block connects the gateway to your IdP and decides who is let in.

KeyDefaultNotes
issuerrequiredDiscovery base; must serve /.well-known/openid-configuration. Use HTTPS in production. A loopback issuer is blocked by the SSRF guard unless the gateway environment sets CLAUDE_GATEWAY_ALLOW_LOOPBACK=1.
client_idrequiredFrom the OAuth client registration.
client_secretrequired unless using private_key_jwtFrom the registration. Omit with certificate authentication.
allowed_email_domainsnoneReject tokens whose email is outside these domains (case-insensitive). Tokens with email_verified: false are always rejected regardless.
allowed_groupsnoneOnly members of these groups may sign in. Exact, case-sensitive match against groups_claim. Nested groups are not expanded.
groups_claimgroupsClaim holding group membership. Entra app roles arrive as roles. Accepts a JSON Pointer such as /realm_access/roles for nested claims.
google_groupsnoneFor Google Workspace, whose ID tokens carry no groups. Set service_account_json_path (key with domain-wide delegation on https://www.googleapis.com/auth/admin.directory.group.readonly) and admin_email (a real Workspace admin to impersonate). Group email addresses become the user's groups.
email_claimemailClaim holding the email. A key, JSON Pointer or list of fallbacks. ADFS and Entra B2C often need upn or preferred_username.
scopes[openid, profile, email, offline_access]Full override. Must include openid. Dropping offline_access removes refresh tokens, so developers re-authenticate every ttl_hours.
scope_on_refreshfalseResend the scope list on refresh. Needed for IdPs (Okta documents this) that only return an ID token on refresh when asked for openid. Pair with userinfo_fallback if groups go missing on refresh. Unset it if refreshes start failing with invalid_scope. Server v2.1.260+.
extra_auth_paramsnoneExtra query parameters on the authorisation request, such as access_type: offline, domain_hint or acr_values. Cannot override state, nonce, redirect_uri, PKCE, scope, response_type, response_mode or client_id.
userinfo_fallbackfalseFill missing email or groups from /userinfo. Needed for Keycloak lightweight tokens, the Okta org server and minimal ADFS tokens. The ID token remains authoritative.
use_pkcetrueSend an S256 PKCE challenge.
clock_skew_seconds0Tolerance for clock drift when checking token times.
token_endpoint_auth_methodautoclient_secret_basic, client_secret_post or private_key_jwt.
client_assertionnoneprivate_key_pem and certificate_pem for private_key_jwt. Server v2.1.284+.
id_token_signed_response_algRS256Set for IdPs signing with ES256, PS256 or EdDSA.
additional_authorized_partiesnoneExtra accepted azp values, for Keycloak brokering and token exchange.
discovery_urlderivedFetch discovery from here instead; path must contain /.well-known/.
use_proxyunsetRoute IdP calls through HTTPS_PROXY/HTTP_PROXY. Server v2.1.227+.
form_action_originsnoneExtra origins for the /device page's CSP form-action, for IdPs that redirect via a second host (Entra to ADFS, hub-and-spoke Okta, SSO interceptors).
ca_cert_pemsystem storePEM content of a CA to trust for IdP calls only. Load a file with ${file:...}.

A Keycloak realm behind internal PKI, gating on a nested role claim, might look like this:

oidc:
  issuer: https://sso.northwind.internal/realms/engineering
  client_id: claude-gateway
  client_secret: ${KEYCLOAK_CLIENT_SECRET}
  groups_claim: /resource_access/claude-gateway/roles
  allowed_groups: [claude-users]
  userinfo_fallback: true
  ca_cert_pem: ${file:/etc/gateway/northwind-root-ca.pem}

Certificate client authentication

Some IdPs, Microsoft Entra in particular, can authenticate the OAuth client with a certificate instead of a secret. Set token_endpoint_auth_method: private_key_jwt (server v2.1.284+). The gateway then signs a short-lived RS256 JWT with the private key at every sign-in and refresh, identifying the certificate by x5t and x5t#S256 thumbprint headers rather than a kid.

  1. Create an unencrypted RSA key of at least 2048 bits (PKCS#8 or PKCS#1 PEM) and a certificate for it. For example:

    openssl req -x509 -newkey rsa:3072 -nodes -days 730 \
      -subj "/CN=northwind-claude-gw" \
      -keyout gw-oidc.key -out gw-oidc.crt
    
  2. Upload the certificate (never the key) to the app registration.

  3. Reference both files, and remove client_secret (setting both refuses to boot):

    oidc:
      issuer: https://login.microsoftonline.com/<tenant-id>/v2.0
      client_id: <application-id>
      token_endpoint_auth_method: private_key_jwt
      client_assertion:
        private_key_pem: ${file:/run/secrets/gw-oidc.key}
        certificate_pem: ${file:/run/secrets/gw-oidc.crt}
    

    certificate_pem must be a single certificate, without its chain, whose public key matches the key.

  4. Restart and look for an oidc: client authentication private_key_jwt line in the boot log showing the CN, SHA-1 thumbprint and expiry. Compare the thumbprint with the IdP. An expired or not-yet-valid certificate only produces a warning at boot, so test with a real sign-in.

To rotate, upload the new certificate alongside the old one, swap the files and restart (rolling restarts are fine because the IdP accepts both), then remove the old certificate once every replica is on the new one.

Outbound proxies and the IdP

Inference upstreams always honour HTTPS_PROXY and HTTP_PROXY. Calls to the IdP (discovery, JWKS, token, userinfo) go direct unless oidc.use_proxy: true. If a proxy variable is set, use_proxy is unset and the issuer is not in NO_PROXY, the gateway stays direct and logs a boot notice asking you to decide; use_proxy: false silences it.

With use_proxy: true, the gateway resolves each IdP hostname itself and asks the proxy to CONNECT to the IP, so the proxy must allow CONNECT to the IP of every host the discovery document lists. Use an http:// proxy URL. ca_cert_pem and the SSRF guard still apply.

Proxy-only egress

If the gateway can reach the outside world only through a forward proxy and cannot resolve public DNS, or the proxy refuses CONNECT to raw IPs, set the environment variable CLAUDE_GATEWAY_PROXY_IS_EGRESS_BOUNDARY=1 next to HTTPS_PROXY (server v2.1.277+). It is deliberately an environment variable rather than YAML so the config file cannot weaken the address check.

HTTPS_PROXY=http://egress.northwind.internal:3128
NO_PROXY=
no_proxy=
CLAUDE_GATEWAY_PROXY_IS_EGRESS_BOUNDARY=1

It only activates when a proxy variable is set, both NO_PROXY and no_proxy are empty, and CLAUDE_GATEWAY_ALLOW_LOOPBACK is off. Otherwise the gateway logs which condition failed and keeps default behaviour. When active, it logs one network: line at boot.

Outbound trafficDefault with HTTPS_PROXYWith proxy-only egress
Anthropic API upstreams, Workload Identity Federation exchange, telemetry exportsResolved and checked locally, then CONNECT to the IP (collectors in NO_PROXY go direct)Hostname handed to the proxy
IdP callsDirect, unless use_proxy: trueHostname handed to the proxy, unless use_proxy: false
Bedrock, Claude Platform on AWS, Agent Platform, Foundry, Google group lookupsHostname handed to the proxySame

Warning: In this mode the gateway no longer catches hostnames that resolve to dangerous addresses. Only enable it if the proxy itself refuses cloud metadata endpoints such as 169.254.169.254 and metadata.google.internal, link-local ranges and its own loopback, by resolved address and not just by name.

session

KeyDefaultNotes
jwt_secretrequiredAt least 32 bytes of entropy (openssl rand -base64 32). Signs HS256 bearer tokens. Can be a list: the first entry signs, all entries verify. To rotate, prepend the new secret, wait ttl_hours, then remove the old one.
ttl_hours1Token lifetime. Clients refresh silently before expiry if the IdP issues refresh tokens. Shorter means faster deprovisioning. If your IdP cannot grant offline_access, raise it to 8 or 12 so people are not sent back to the browser hourly.

store

KeyDefaultNotes
postgres_urlrequiredpostgres:// or postgresql:// with a single host. The role must be able to create and alter tables, since migrations run at boot and on upgrade.
usernamefrom URLOverrides the URL's user.
passwordfrom URLKeeps the credential out of the URL. Any characters allowed; wins over URL credentials.
max_connections5Pool size per replica. Raise for a dedicated database when spend limits are on, keeping replicas × pool below the server's own limit.
connect_timeout_seconds51 to 60. Server v2.1.274+; earlier versions refuse to boot with it set.
readiness_grace_seconds00 to 3600. How long /readyz stays ready after Postgres stops answering. Server v2.1.282+.

For a local throwaway database, a trust-auth container is enough: docker run --rm -p 5432:5432 -e POSTGRES_HOST_AUTH_METHOD=trust postgres.

upstreams

upstreams is an ordered list. Each request goes to the first upstream able to serve the requested model.

The gateway fails over to the next entry on 5xx, 429, 401, 403, 404, 501 or a timeout. Other 4xx codes are treated as the request's fault and returned immediately. 404 failover needs server v2.1.198+. If two entries use the same provider, each needs a distinct name:.

Cloud SDK clients (Bedrock, Claude Platform on AWS, Agent Platform, Foundry) are built at startup and refresh their own credentials, so rotating cloud credentials needs no restart. Static Anthropic keys and bearers are read once at startup.

Errors the developer sees

  • If an upstream returns a status that does not trigger failover, its response goes back and no further upstreams are tried.
  • If every attempt failed over, the gateway returns the last 429; failing that the last 401 or 403, then the last 404, then the last 501; failing all of those, its own 502 all upstreams failed (N attempted), where N counts every entry including ones skipped for not serving the model.

Status codes are preserved. An Anthropic API upstream's error body passes through untouched. Cloud upstreams can mention account IDs, role ARNs and project IDs, so the gateway logs their full text but shows developers:

  • For 400/413 in Anthropic's error envelope (Claude Platform on AWS, Agent Platform and Foundry use it), the upstream's own message, such as prompt is too long.
  • For 400/413 in the provider's own format, a capability_rejected: token, or upstream rejected the request / request too large for this upstream if unclassified.
  • For anything else, generic text such as upstream rate limit exceeded.

Bedrock's "input too long" error becomes capability_rejected: prompt_too_long, which Claude Code treats as a cue to compact, as described in errors. This behaviour needs server v2.1.233+.

provider: anthropic

upstreams:
  - provider: anthropic
    auth:
      api_key: ${ANTHROPIC_KEY_PLATFORM_TEAM}
    # base_url: https://api.anthropic.com    (default)

auth.api_key sends x-api-key. auth.oauth_token sends Authorization: Bearer instead, for organisations issuing short-lived tokens; it is read once at startup, so remount and restart to refresh.

To avoid static credentials entirely, use Workload Identity Federation: create a federation rule in the Claude Console, mount your workload's OIDC token as a file, and the gateway exchanges it for short-lived bearers, re-reading the file on every exchange.

upstreams:
  - provider: anthropic
    auth:
      federation_rule_id: ${CLAUDE_FED_RULE}
      organization_id: ${CLAUDE_ORG_ID}
      identity_token_file: /var/run/secrets/tokens/claude-audience
      # workspace_id: wrkspc_...      needed when the rule spans several workspaces
      # service_account_id: svac_...  optional target check
Per-user identity headers for a proxy you run

If base_url points at a proxy you operate (LiteLLM, for example), set forward_user_identity: true (server v2.1.233+) and the gateway adds:

HeaderValue
x-litellm-end-user-idDeveloper's email, when known
x-claude-gateway-user-idThe token's sub
x-claude-gateway-user-emailDeveloper's email, when known

A 429 from that proxy on a request carrying an email is treated as a per-user denial and returned to the developer rather than failed over (from server v2.1.267). Requests without an email omit the email headers, and their 429s fail over as normal. The gateway refuses to boot if you enable this with base_url left at the Anthropic API.

provider: bedrock

upstreams:
  - provider: bedrock
    region: eu-west-2
    auth: {}
    # auth: { aws_access_key_id: ..., aws_secret_access_key: ..., aws_session_token: ... }
    # auth: { aws_bearer_token: ${BEDROCK_API_TOKEN} }
    # base_url: https://bedrock-runtime-fips.us-east-1.amazonaws.com

An empty auth uses the AWS default chain (env vars, ~/.aws/credentials, ECS task role, EC2 metadata, IRSA). Explicit keys must come as a complete pair; a session token without both keys fails boot.

What the AWS side needs:

  • IAM: bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream on the inference-profile ARNs (for the built-in US catalogue, arn:aws:bedrock:<region>:<account>:inference-profile/us.anthropic.*) and on arn:aws:bedrock:*::foundation-model/anthropic.*. Add bedrock:CountTokens on the foundation-model ARNs so the gateway can count abandoned requests for spend limits for free; without it, it sends a one-token request instead.
  • Model access: on by default in commercial regions, apart from Anthropic's one-time use case form in the Bedrock console. See Amazon Bedrock.
  • EKS: an IRSA role annotated on the service account as eks.amazonaws.com/role-arn. ECS/EC2: attach the role to the task or instance profile.
  • Region: the endpoint region. Cross-region inference profiles still route across their geography. Non-US regions and provisioned throughput need a models: block.
Apply an Amazon Bedrock guardrail
    guardrail:
      id: gr-7hq2northwind      # ID or full ARN
      version: "3"              # quoted; or DRAFT

Server v2.1.281+. Grant bedrock:ApplyGuardrail to whichever principal signs the requests. Either every bedrock upstream sets a guardrail or none does, and a mantle upstream cannot coexist with a guardrail. Other providers in the list are not covered. A request already carrying an amazon-bedrock-* body field is rejected with 400. Guardrail input tags are not added, so filters that rely on tagged input do not run.

Bedrock in another AWS account

assume_role (server v2.1.281+) makes the gateway use its own identity only to call sts:AssumeRole, then sign Bedrock calls with the one-hour credentials returned.

upstreams:
  - name: bedrock-research
    provider: bedrock
    region: eu-west-2
    auth: {}
    assume_role:
      role_arn: arn:aws:iam::444455556666:role/claude-gw-invoke
      external_id: ${RESEARCH_EXTERNAL_ID}
KeyPurpose
role_arnarn:aws:iam:: or arn:aws-us-gov:iam:: role carrying the Bedrock permissions (plus bedrock:ApplyGuardrail if needed).
external_idSent on every assume call. Quote it if numeric.
session_nameemail or sub for per-developer sessions. Unset means one shared session called claude-apps-gateway.

The target role's trust policy names the gateway's own principal and allows sts:AssumeRole, with an sts:ExternalId condition if you use one. If STS fails, the gateway does not fall back to its own credentials for that upstream; it logs the error and moves to the next entry. It calls the regional sts.<region>.amazonaws.com endpoint; for FIPS set AWS_USE_FIPS_ENDPOINT=true. assume_role cannot be combined with aws_bearer_token.

To keep a model served only from this account, give it a custom model ID whose upstream_model map lists only this upstream. Built-in model names can still reach it in order, so list it last unless it should serve those too.

Per-developer AWS cost attribution

Add session_name: email (or sub) and the gateway assumes the role once per developer per hour per replica, naming the session after them. AWS then attributes each person's calls to their own assumed-role session. Characters outside ASCII letters, digits and _+,.@- are encoded as =XX, and names over 64 characters are shortened with a hash. Developers whose token lacks the claim are not routed through this upstream. The token-count fallback for abandoned requests still runs under the shared claude-apps-gateway session. Set this on every Bedrock upstream for strict attribution; the AWS walkthrough is in Claude apps gateway on AWS.

provider: mantle

The Bedrock Mantle endpoint (server v2.1.283+):

upstreams:
  - provider: mantle
    region: us-east-1
    models: [claude-sonnet-4-6, claude-haiku-4-5]
    auth: {}
  - provider: bedrock
    region: us-east-1
    auth: {}
KeyRequiredNotes
regionYesEndpoint becomes https://bedrock-mantle.<region>.api.aws/anthropic.
modelsYesModels your account has on Mantle, as clients name them. Everything else skips to the next upstream.
authNoSame keys and rules as Bedrock.
base_urlNoOverride; keep the trailing /anthropic.

Grant Mantle's own IAM actions (listed on Amazon Bedrock). Mantle does not take assume_role or guardrails. For a Mantle ID the gateway does not know, define it in models: and add that id to the upstream's models list too.

provider: anthropicAws

Claude Platform on AWS (server v2.1.198+) serves the first-party API at aws-external-anthropic.<region>.api.aws, with first-party model IDs, beta headers honoured and count_tokens available.

upstreams:
  - provider: anthropicAws
    region: eu-west-1
    workspace_id: wrkspc_01northwind
    auth:
      api_key: ${CPA_API_KEY}
    # or auth: {} for SigV4 via the default chain
KeyNotes
regionRequired. Lowercase letters, digits and hyphens.
workspace_idRequired. Sent as a header on every request.
auth.api_keySent as x-api-key; wins if SigV4 credentials are also present.
auth.aws_access_key_id / auth.aws_secret_access_keyExplicit SigV4, both or neither, with optional aws_session_token.
base_urlOverride the endpoint.

It signs for the aws-external-anthropic service in its own account, so a Bedrock IAM role does not authorise it. The built-in catalogue routes to it with no models: block; key curated entries as anthropicAws:. See Claude Platform on AWS.

provider: vertex

Google Cloud's Agent Platform:

upstreams:
  - provider: vertex
    region: europe-west1
    project_id: northwind-ai-prod
    auth: {}
    # auth: { service_account_json: /secrets/vertex-sa.json }

Empty auth uses Application Default Credentials. service_account_json takes a file path, not contents. region: global uses the global endpoint so Google picks an available region for you. Grant roles/aiplatform.user (or a role with aiplatform.endpoints.predict), enable aiplatform.googleapis.com, and enable the Claude models in Model Garden. On GKE, bind via Workload Identity with the iam.gke.io/gcp-service-account annotation. More in Google Cloud and Claude apps gateway on Google Cloud.

provider: foundry

upstreams:
  - provider: foundry
    resource: northwind-foundry-uks
    auth: { use_azure_ad: true }
    # auth: { api_key: "${FOUNDRY_KEY}" }

use_azure_ad: true uses DefaultAzureCredential (managed identity, workload identity on AKS, Azure CLI, environment). The endpoint derives from resource; set base_url for sovereign clouds. Grant Azure AI User or Cognitive Services User on the resource. Foundry uses your own deployment names, so a models: block is required. Quote ${...} inside inline { } maps. See Microsoft Foundry.

Static headers on upstream requests

headers: (server v2.1.277+) adds fixed headers to one upstream's requests, useful when a proxy in front of the provider routes by header.

  - provider: foundry
    resource: northwind-foundry-uks
    base_url: https://ai-egress.northwind.internal
    auth: { use_azure_ad: true }
    headers:
      x-cost-centre: "4410"
      x-egress-token: ${EGRESS_TOKEN}

Values must be printable ASCII without leading or trailing spaces; quote numbers and booleans. An empty ${VAR} stops boot. Headers go on /v1/messages and count_tokens calls, including ones that failed over to this upstream, but not on Bedrock's abandoned-request CountTokens call or the federation token exchange. On SigV4 upstreams they are part of the signature, so proxies must pass them unchanged. Reserved names fail boot: authorization, x-api-key, host, content-type, user-agent, and anything starting anthropic-, x-goog-, x-amz- or x-amzn-.

Multiple upstreams

Every request starts at the top of the list. The gateway does not remember failures, so a dead upstream is retried by every request that reaches it. For Anthropic API upstreams, timeouts.upstream_ttfb_ms limits that wait; other providers can wait up to an hour for a first byte. Upstreams that cannot resolve the model are skipped without a network call.

A UK-first layout with overflow and a last-resort fallback:

upstreams:
  - name: bedrock-ldn-pt
    provider: bedrock
    region: eu-west-2
    auth: {}
  - name: bedrock-eu-od
    provider: bedrock
    region: eu-central-1
    auth: {}
  - name: vertex-eu
    provider: vertex
    region: europe-west1
    project_id: northwind-ai-prod
    auth: {}

models:
  - id: claude-sonnet-4-6
    label: Claude Sonnet 4.6
    upstream_model:
      bedrock-ldn-pt: arn:aws:bedrock:eu-west-2:111122223333:provisioned-model/9z8y7x
      bedrock-eu-od: eu.anthropic.claude-sonnet-4-6
      vertex-eu: claude-sonnet-4-6
GoalHow
Several regionsOne upstream per region. Use models: for region-pinned IDs.
Several accountsOne upstream per account, using assume_role or explicit credentials.
Provisioned throughput firstMap the PT ARN on that upstream only, so it is exhausted (429) before overflow.
VPC or FIPS endpointsbase_url on the upstream.
Restrict a model to some upstreamsOnly works for a custom model ID: upstreams missing from its map are skipped. For built-in models the map only changes which ID is sent. mantle upstreams are tried only for their own models list.

Note: Failing over between providers, or to the Anthropic API, changes which agreement and geography govern that request. Think about this before adding a cross-provider fallback.

Optional sections

admin

Enables the spend-limits admin API and live enforcement. Usage is described in spend limits.

KeyDefaultNotes
write_keysnoneList of {id, key}. Full access. Keys at least 32 characters; IDs unique across both lists.
read_keysnoneList of {id, key}. Every GET endpoint.
admin_groupsnoneIdP groups whose members get full access via their normal gateway token. Audits as oidc:<sub>. Empty entries fail boot.
blocked_messagenoneAppended verbatim to the blocked developer's 429.
audit_retention_days365Age of admin_audit rows before sweeping.
spend_retention_months13Age of spend counters before sweeping.
identity_retention_days90Time since last activity before principal_emails (personal data) rows are swept.
group_limit_modeminmin applies the strictest group cap, max the most generous.

enforcement

KeyDefaultNotes
fail_closed_on_errorfalsetrue blocks requests when Postgres is unreachable. Requires admin; the gateway refuses to boot otherwise.

pricing

Replaces list prices in the spend meter, and in the figures developers see, with your contracted rates. Still USD, still an estimate. Requires server v2.1.227+ and either an admin block or (v2.1.268+) a managed block with at least one policy.

pricing:
  multiplier: 0.9
  overrides:
    - upstream: bedrock-ldn-pt
      model: claude-sonnet-4-6
      input: 2.70
      output: 13.50
      cache_read: 0.27
      cache_write: 3.375
KeyNotes
multiplierDefault 1. Applied to every amount. Greater than 0, at most 10.
overridesRows of upstream, model, input, output, cache_read, cache_write, in USD per million tokens. All four rates required, each above 0 and at most 10000.

Matching rules:

  • A row applies to requests that the named upstream serves for that model, including fast mode requests, which then meter at the same rates.
  • A built-in ID like claude-sonnet-4-6 covers all its dated, regional Bedrock and Agent Platform forms. Any other string is matched case-insensitively against the client's ID or the upstream string.
  • When rows overlap, the most specific wins: exact upstream string, then exact client ID, then built-in model name.
  • Unknown upstream names, or two rows for the same upstream and model, fail boot. Unusable rows produce a warning.
  • Web search stays at $0.01 per request, multiplied.

For per-region rates, give each region its own named upstream.

Mark prices up

From server v2.1.271 a multiplier above 1 (up to 10) meters more than list, for internal chargeback. With admin configured, caps are reached sooner and the gateway warns at boot. It never changes what the provider bills. Clients need v2.1.271+ to display the markup; older ones ignore it. Older servers refuse to boot with a multiplier above 1.

Send the rates to signed-in clients

From server v2.1.268, the gateway injects the pricing into each managed policy as the modelPricing setting, so /usage, the status line and OpenTelemetry show your rates (clients v2.1.242+). It adds the multiplier and the override row of the first upstream serving each model ID. Developers matched by no policy see list price. To opt a policy out, set modelPricing: {} in its cli block; a policy that sets its own modelPricing keeps it untouched.

models

A curated list served at /v1/models and used to translate IDs per upstream. Required for non-US Bedrock regions, provisioned throughput ARNs and Foundry deployment names.

auto_include_builtin_models: false   # true keeps the built-in catalogue as well
models:
  - id: claude-opus-4-8
    label: Claude Opus 4.8
    description: For architecture work and hard debugging
    upstream_model:
      bedrock-eu-od: eu.anthropic.claude-opus-4-8
      northwind-foundry-uks: opus-prod-deployment

Each key under upstream_model must equal an upstream's name (which defaults to the provider name). A key matching no upstream fails boot.

managed

Policies matched by IdP group or email domain, served per user at GET /managed/settings with ETag caching. Policies are checked in order; the first match wins and is layered on top of the match: {} catch-all.

managed:
  policies:
    - match: { groups: [contractors-external] }
      cli:
        availableModels: [claude-sonnet-4-6]
        permissions:
          deny: ["WebFetch", "Bash(curl *)"]
    - match: { email_domain: northwind-labs.co.uk }
      cli:
        availableModels: [claude-opus-4-8, claude-sonnet-4-6]
    - match: {}
      cli:
        availableModels: [claude-opus-4-8, claude-sonnet-4-6, claude-haiku-4-5]
        enforceAvailableModels: true
        permissions:
          deny: ["Read(./.env*)"]

How keys combine with the catch-all base:

Key typeExamplesBehaviour
Allow listsavailableModels, permissions.allowThe specific policy's list replaces the base's.
Deny lists and hookspermissions.deny, permissions.ask, disabledMcpjsonServers, deniedMcpServers, blockedMarketplaces, every hooks event arrayUnion of both, so org-wide denials and audit hooks cannot be dropped.
Recordsenv, modelOverrides, skillOverridesShallow merge: the policy overrides the keys it sets.
MatcherMatches
{}Everyone who authenticated.
{ groups: [a, b] }Any listed group in the token's groups claim, case-sensitive.
{ email_domain: example.com }Text after the last @ in the email, case-insensitive. One domain per policy.
{ groups: [a], email_domain: example.com }Both conditions.

A user matching nothing gets every catalogue model and no managed settings, so a catch-all at the end is strongly advised.

availableModels is also enforced at /v1/messages, returning 400 for anything else; an empty list denies everything. The gateway rejects a missing or empty model with model is required (server v2.1.228+) and a non-string with model must be a string (v2.1.221+).

The gateway keeps no user directory: it reads groups from each token. There is nothing to pre-create and no SCIM endpoint; run lifecycle management in your IdP. Policy edits reach connected CLIs on the next hourly poll (apart from launch-only settings). Group membership changes take effect at the next token refresh, bounded by ttl_hours.

Start sessions on a model the policy allows

Gateway sessions default to the Opus model that the opus alias resolves to. If a policy's availableModels excludes it, new sessions get 400s until the developer runs /model. Set enforceAvailableModels: true in the same cli block. If the list contains an alias or a built-in ID, the default then resolves inside it. If the list contains only custom IDs, also set model:

    - match: { groups: [secure-enclave] }
      cli:
        availableModels: [opus-isolated]
        enforceAvailableModels: true
        model: opus-isolated

See model configuration for both keys.

Matcher values that stop the gateway at boot

From server v2.1.232, boot fails on an empty groups list, an empty string in groups or admin_groups, an empty email_domain, or an email_domain containing @, whitespace or a comma (after trimming and stripping one leading @). Earlier versions accepted these with surprising results: an empty domain with no groups matched everyone, and an empty admin_groups entry could grant admin to a user whose token also carried an empty group.

What goes in cli

Each cli value is a full managed-settings.json document written as YAML, applied by the CLI at the managed tier in place of server-managed settings. Keys restricted to OS-level sources, such as policyHelper and wslInheritsWindowsSettings, are ignored. The gateway validates against the settings schema of its own version, so newer top-level keys need a gateway upgrade first. Open-ended areas such as env, pluginConfigs and nested permissions keys accept anything. The older name settings still works as an alias.

The keys I reach for first:

KeyEnforced byEffect
availableModelsGateway and CLIModel allowlist, also checked server-side.
permissions.allow / .denyCLITool rules; see permissions.
permissions.disableBypassPermissionsMode: disableCLIBlocks bypass mode and --dangerously-skip-permissions.
allowManagedPermissionRulesOnlyCLIOnly managed permission rules count.
envCLIVariables merged into the CLI process.
hooksCLIOrg-wide hooks. Commands run on developer machines, so paths must exist on every OS the policy covers.
managedMcpServersCLIRemote (http, sse) MCP servers given to everyone matched. Gateway and clients v2.1.259+.

Because these arrive over the network, interactive clients show a security approval dialog before applying hooks, shell-executing settings like apiKeyHelper and statusLine, sandbox binary paths (sandbox.bwrapPath, sandbox.socatPath, sandbox.ripgrep), sandbox settings that intercept or weaken isolation such as sandbox.network.tlsTerminate, and sensitive env values. Non-empty proxy, base-URL and OTEL_EXPORTER_OTLP_ENDPOINT values always need approval; model selection and numeric limits do not. Declining exits the session. Non-interactive runs (claude -p, the Agent SDK) apply the settings for that run without recording approval. Details are in server-managed settings.

Since the telemetry section pushes OTEL_EXPORTER_OTLP_ENDPOINT, enabling forward_to triggers that dialog once for every interactive client.

Context window in terminal sessions

Gateway terminal sessions use the 1M context window for Opus 4.7+, Sonnet 5+ and Fable models without any [1m] suffix, compacting near 967K tokens (client v2.1.287+). To compact at 200K instead, push CLAUDE_CODE_AUTO_COMPACT_WINDOW: "200000" in a policy's env (no approval needed). To disable 1M entirely, push CLAUDE_CODE_DISABLE_1M_CONTEXT: "1", which developers approve interactively. See context window.

MCP servers in a policy

Use managedMcpServers in cli, not the .mcp.json key mcpServers, which is rejected at boot. Entries are checked with the same rules the client applies, as described in managed MCP. ${VAR} references are expanded on the gateway before delivery, so every matched client receives the literal value.

Claude Desktop overlay

The same gateway can serve Claude Desktop. Set Desktop's bootstrapUrl to <public_url>/user/bootstrap. Desktop derives the OAuth issuer from it, runs the device sign-in, and pulls its configuration from the response. This needs server v2.1.203+ and an explicit opt-in: /user/bootstrap returns 404 unless the matching policy has a desktop key (even desktop: {}, and a key on the catch-all opts in everyone who inherits it). Requests are audited as desktop_bootstrap.serve or desktop_bootstrap.denied.

What the gateway derives for Desktop from the policy:

  • Models, from availableModels.
  • Disabled tools, from bare tool names in permissions.deny. A disabledBuiltinTools value in desktop is unioned in.
  • Egress allowlist, from sandbox.network.allowedDomains, unless desktop.coworkEgressAllowedHosts replaces it.
  • An OTLP endpoint pointing at the gateway plus identity attributes, when telemetry.forward_to and public_url are both set. Desktop exports http/protobuf, or http/json if the policy's env sets that protocol.

Keys with no Desktop equivalent, like hooks and scoped rules such as Bash(npm *), are left out.

To set Desktop keys directly, add a desktop block next to cli using flat key names from Desktop's managed configuration:

    - match: { groups: [design-team] }
      cli:
        availableModels: [claude-sonnet-4-6]
        enforceAvailableModels: true
      desktop:
        chatTabEnabled: true
        isLocalDevMcpEnabled: false
        banner: { text: "Northwind design workspace" }

At boot the gateway validates desktop against Desktop's own schema and fails on unknown keys, invalid values, keys the gateway computes itself (inference connection, model list, OTLP relay), MDM-only keys such as bootstrapUrl, and legacy aliases. Deprecated shapes only warn. disabledBuiltinTools, coworkEgressAllowedHosts and Desktop's array-form managedMcpServers need server v2.1.232+. Newer Desktop keys need a newer gateway; for example userPluginMarketplacesEnabled and userPluginUploadsEnabled need v2.1.260+ and Desktop 1.37937.0+, and blockReadsOutsideWorkingDirectories, disableBypassPermissionsMode, configRecheckIntervalMinutes and sshClientPath need v2.1.281+. orgPluginSettings is served in the array form read by Desktop 1.15200.0+.

Inheritance from the base desktop block follows the same idea as cli, with two protective exceptions: disabledBuiltinTools is unioned, and a non-allow value in the base's builtinToolPolicy cannot be relaxed by a role policy. Arrays and nested objects such as banner are replaced whole.

Policy changes reach Desktop at next launch. A running Desktop checks every 10 minutes, applies a few settings live, and shows a Relaunch Claude Desktop card for the rest, forcing a restart after 24 hours (tunable with relaunchEnforcementHours, server v2.1.260+, Desktop 1.40609.0+; 0 prompts immediately).

Extended context in Claude Desktop

From server v2.1.284, Desktop's picker offers a 1M option for listed models that support it (Opus 4.6+, Sonnet 4.6+, Fable), unless any upstream that could serve the entry maps it to a model without 1M support, or the entry names no Claude model at all. modelPrefer1mContext: true in desktop starts new users on the 1M variant. For older servers or custom aliases, list the model twice in models, once with [1m] appended and the same upstream_model map; only do this for models your upstreams really serve at 1M. To remove the option, push CLAUDE_CODE_DISABLE_1M_CONTEXT: "1" in cli.env and delete any [1m] entries.

Precedence with other managed sources

Gateway-delivered settings outrank an MDM profile or local managed-settings.json on the same device, with the exceptions listed under precedence in managed settings. A policyHelper runs only when the gateway delivers nothing. Gateway policy applies to every Claude Code process on the machine, and signed-in sessions refuse to start if the gateway is unreachable.

telemetry

The CLI sends OTLP/HTTP exports to the gateway, which relays them verbatim to each destination. Exports from /login sessions are stamped with user.id, user.email and user.groups from the gateway token, so per-person reporting needs nothing on the developer side. Desktop and Cowork exports carry user.email, user.groups (server v2.1.265+, Desktop 1.24012+) and enduser.sub (server v2.1.274+); their user.id is anonymous, so join terminal user.id against Desktop enduser.sub. Over-long or awkwardly formatted group lists and subjects are omitted from Desktop telemetry rather than truncated. None of this goes to Anthropic.

telemetry:
  forward_to:
    - url: https://otel.northwind.internal:4318
      headers:
        Authorization: Bearer ${OTEL_INGEST_TOKEN}
      metrics: true
      logs: true
      traces: false
    - url: https://otlp.eu.grafana.example.net/otlp
      headers:
        Authorization: ${GRAFANA_BASIC_AUTH}
  resource_attributes:
    service.namespace: claude
    deployment.environment.name: production

Warning: Each destination defaults to metrics only. Logs and traces can contain full shell commands, tool inputs and file paths. Turn them on only for destinations with suitable access control and retention.

URLs must be https://. http://localhost:<port> validates but is blocked at export time, and http://127.0.0.1 or http://[::1] fail boot, unless CLAUDE_GATEWAY_ALLOW_LOOPBACK=1 is set. Exports use HTTPS_PROXY when set; to reach an internal collector directly add it to NO_PROXY by exact name or by leading-dot domain (server v2.1.277+). CIDRs do not match.

With forward_to and public_url both set, the gateway switches telemetry on for clients by pushing CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER, OTEL_LOGS_EXPORTER and OTEL_TRACES_EXPORTER (each otlp if any destination wants that signal, otherwise none), OTEL_EXPORTER_OTLP_ENDPOINT=<public_url> and OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf, plus OTEL_RESOURCE_ATTRIBUTES when you define labels. These override anything set locally, and signed-in clients ignore locally configured endpoints. Signals with no destination are discarded.

Traces also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 pushed through a policy's env; set 0 in a group's policy to stop that group tracing. The metrics and events themselves are documented in monitoring usage.

Add your own labels

telemetry.resource_attributes (server v2.1.281+) attaches fixed OpenTelemetry resource attributes to gateway sessions. Rules, enforced at boot: names use letters, digits, ., _ and -; names starting user., enduser. or identity., and service.name, service.version, claude.deployment_mode, host.arch, os.type, os.version and wsl.version are reserved; values are non-empty printable ASCII without spaces or , ; = \ " %, at most 255 characters after percent-encoding, and quoted if they look like numbers or booleans. A policy that sets OTEL_RESOURCE_ATTRIBUTES in env overrides the labels for terminal sessions. Labels are copied onto metric data points too; see monitoring usage for cardinality controls.

Export directly to your collector

To skip the relay, set OTEL_EXPORTER_OTLP_ENDPOINT to your collector's https:// base URL in a policy's env, with OTEL_EXPORTER_OTLP_HEADERS for auth (clients v2.1.265+). Claude Code appends /v1/metrics, /v1/logs or /v1/traces and never sends the gateway token there. Developers approve the endpoint interactively. The client falls back to the relay if the value did not come from the gateway, is not HTTPS (or loopback HTTP), does not resolve to a /v1/<signal> path without query or fragment, points at the gateway host, or if any otelHeadersHelper is configured. If the gateway is not already pushing the telemetry variables, also set CLAUDE_CODE_ENABLE_TELEMETRY=1, the exporter selectors and the protocol. Exports stop when the developer signs out or switches gateway.

When a destination fails

There is no buffering or retry. A failed export is dropped and logged; the client always gets success. After five consecutive failures to one destination, forwarding pauses in 30 second stretches until one succeeds. Responses 400, 413, 415, 422 and 431 mean the payload was refused; they neither count towards nor reset the failure streak, and are logged on the first occurrence and every hundredth.

HTTP tuning

BlockKeyDefaultNotes
access_controlallow_cidrs, deny_cidrsemptyClient IP filter after trusted_proxies resolution. Deny is checked first. A non-empty allow list makes the gateway default-deny. /healthz and /readyz are exempt from the allow list. An unparseable forwarded address is refused with 403 (xff_unparseable) where a list applies.
limitsmax_request_bytes32 MiBLarger bodies get 413 before buffering.
limitsmax_request_header_bytesunsetLowers the built-in 256 KiB header limit; over the limit returns 431.
limitsmax_url_lengthunsetOver-long URLs return 414.
timeoutsupstream_ttfb_ms120000Wait for response headers from the Anthropic API upstream only.
rate_limitsdevice_authorization.max / .window_seconds30 / 600Per-IP limit on starting a sign-in. Raise for large offices behind one NAT.
rate_limitsdevice_verify.max / .window_seconds10 / 600Per-IP limit on entering codes at /device; this is the brute-force defence.

With both access lists empty the gateway serves any address, and it warns about that at boot (suggesting 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10, 127.0.0.0/8, ::1/128 and fc00::/7) and once at runtime with an access.public_client audit event for the first non-private client. A relay that is not in trusted_proxies hides real addresses from both checks, so configure that first.

access_control:
  allow_cidrs: [10.40.0.0/14, 100.64.0.0/10]
listen:
  trusted_proxies: [10.40.8.0/24]

load_test_mode

Server v2.1.282+. The gateway builds and signs each upstream request as normal, throws it away, and streams back canned filler text that starts by saying it is canned.

load_test_mode:
  enabled: true
  reply_tokens: 1200
  reply_seconds: 15
KeyDefaultNotes
enabledrequiredfalse keeps the block but turns it off.
reply_tokens7501 to 100000.
reply_seconds9.50 to 600; 0 sends at once. Non-streaming replies are always immediate.

An x-load-test-user header with up to seven digits makes each number count as a separate developer. The gateway refuses to start in this mode against a database where anyone has recorded spend, so give it a fresh one. CPU per request reads lower than production because nothing is encrypted to a provider; confirm sizing with a small real pilot. Each inference audit event carries load_test: true.

Warning: Never enable this on a gateway people use. Every reply is filler.

A fuller example

This config is the sort of thing I would hand a client as a starting point: Entra sign-in, a UK Bedrock primary with a Claude Platform on AWS fallback, two groups, spend limits and one telemetry destination. Operational log verbosity comes from the CLAUDE_GATEWAY_LOG_LEVEL environment variable (debug, info, warn, error; default info). debug also prints the claim names in each ID token, which is handy when groups are not matching. Audit events are always emitted.

listen:
  public_url: https://claude.tools.northwind.internal
  trusted_proxies: [10.40.8.0/24]

oidc:
  issuer: https://login.microsoftonline.com/<tenant-id>/v2.0
  client_id: <application-id>
  client_secret: ${ENTRA_CLIENT_SECRET}
  allowed_email_domains: [northwind.co.uk]
  groups_claim: roles

session:
  jwt_secret: ${GW_SIGNING_SECRET}

store:
  postgres_url: ${GW_DATABASE_URL}
  max_connections: 15
  readiness_grace_seconds: 120

admin:
  write_keys:
    - { id: finops-terraform, key: "${SPEND_ADMIN_KEY_TF}" }
  admin_groups: [Claude.Admins]
  blocked_message: "Raise a ticket in the Platform Help portal for more budget."

upstreams:
  - name: bedrock-ldn
    provider: bedrock
    region: eu-west-2
    auth: {}
  - name: cpa-dub
    provider: anthropicAws
    region: eu-west-1
    workspace_id: wrkspc_01northwind
    auth: {}

auto_include_builtin_models: false
models:
  - id: claude-opus-4-8
    label: Claude Opus 4.8
    upstream_model:
      bedrock-ldn: eu.anthropic.claude-opus-4-8
      cpa-dub: claude-opus-4-8
  - id: claude-sonnet-4-6
    label: Claude Sonnet 4.6
    upstream_model:
      bedrock-ldn: eu.anthropic.claude-sonnet-4-6
      cpa-dub: claude-sonnet-4-6

managed:
  policies:
    - match: { groups: [Claude.Contractors] }
      cli:
        availableModels: [claude-sonnet-4-6]
        permissions:
          deny: ["WebFetch"]
    - match: {}
      cli:
        availableModels: [claude-opus-4-8, claude-sonnet-4-6]
        enforceAvailableModels: true
        permissions:
          disableBypassPermissionsMode: disable
          deny: ["Read(./.env*)", "Read(./secrets/**)"]
        env:
          DISABLE_UPDATES: "1"

telemetry:
  forward_to:
    - url: https://otel.northwind.internal:4318
      headers:
        Authorization: Bearer ${OTEL_INGEST_TOKEN}

DISABLE_UPDATES stops both background and manual updates, which suits a team that ships Claude Code versions through its own packaging; DISABLE_AUTOUPDATER would stop only background ones.

Client-side managed settings

None of the above tells a laptop where the gateway is. That has to come from the device's own managed settings, because the gateway cannot push the very keys that point clients at it.

{
  "forceLoginMethod": "gateway",
  "forceLoginGatewayUrl": "https://claude.tools.northwind.internal",
  "parentSettingsBehavior": "merge",
  "allowedProviders": ["gateway"]
}
  • forceLoginMethod and forceLoginGatewayUrl route /login to your gateway.
  • parentSettingsBehavior: "merge" keeps Claude Desktop's egress allowlist reaching its embedded sessions; the reasoning is in Claude apps gateway.
  • allowedProviders: ["gateway"] (v2.1.285+) refuses any session not set up for a Cloud gateway, so nobody bypasses it with CLAUDE_CODE_USE_BEDROCK or their own ANTHROPIC_BASE_URL. Only the gateway named in forceLoginGatewayUrl, or one set as ANTHROPIC_BASE_URL in the file's env, is admitted. Keep this key off the gateway host itself: claude gateway refuses to run where it is set.

Deliver the file through MDM to the per-OS location listed in managed settings. A Windows registry policy or macOS managed-preferences plist replaces the file rather than merging with it, so fleets using those mechanisms must put the keys there instead.

forceLoginGatewayUrl, gatewayInternalNetworks and the "gateway" value of forceLoginMethod are honoured only from a managed source on the machine (file, plist, HKLM or policy helper), never from ~/.claude/settings.json or the gateway payload. Conversely, keep forceLoginMethod and forceLoginOrgUUID out of the gateway's cli payload: Claude Code still reads them from there for its startup credential check, which can lock out developers who keep an Anthropic credential on the machine.

For Desktop, set bootstrapUrl in its own managed configuration to <public_url>/user/bootstrap and opt policies in with a desktop key, as covered in the overlay section above.