Skip to content

Deploying and running Claude apps gateway

Production runbook for Claude apps gateway: IdP registration, network address, container image, Kubernetes and Cloud Run, operations, security and troubleshooting.

The overview gets a gateway running on one machine. This page is what you need to take it to production and keep it there: registering with your identity provider, picking an address, building the image, running it on Kubernetes or Cloud Run, and the day-two work of logs, probes, secret rotation and upgrades. Every gateway.yaml key is documented in the configuration reference.

Warning: Decide the gateway's network address before anything else. At /login, Claude Code refuses any gateway whose hostname resolves to a public IP address, even an unreachable one. A gateway can push settings to laptops, including hooks that run shell commands, so the check stops people signing in to a malicious gateway on the internet. Normally you give the gateway a private address reached over the internal network or VPN. If your internal network uses public IPv4 space you own, you can declare one block containing both the gateway and users' machines. If neither works for you, talk to your Anthropic account team.

I work through it in four stages:

  1. Register with the identity provider
  2. Deploy, making the address, image and platform choices
  3. Operate: logs, probes, outages, rotation, upgrades
  4. Review security for the questionnaire

If something fails along the way, jump to troubleshooting, which is organised by error message.

Identity provider

Register a confidential OAuth/OIDC web application with exactly one redirect URI, https://<gateway>/oauth/callback, and assign it to the people or groups who should have access. The gateway authenticates to the IdP with the client secret, or with an uploaded certificate if your IdP uses certificate credentials.

Any OIDC-compliant IdP works (Okta, Microsoft Entra ID, Google Workspace, Keycloak, Dex, PingFederate and so on) provided it:

  • serves /.well-known/openid-configuration, over HTTPS in production. An http:// issuer is accepted, and a loopback issuer additionally needs CLAUDE_GATEWAY_ALLOW_LOOPBACK=1;
  • supports the authorisation-code flow. PKCE is on by default; switch it off with oidc.use_pkce: false if the IdP cannot do it;
  • returns email (and optionally groups) in the id_token, or from the userinfo endpoint when oidc.userinfo_fallback: true is set.

For a private PKI, set oidc.ca_cert_pem.

Provider quirks

IdPWhat to know
OktaThe org authorisation server (https://yourco.okta.com) issues a thin id_token without email or groups, so set oidc.userinfo_fallback: true when it is your issuer. A custom authorisation server such as https://yourco.okta.com/oauth2/default can include them directly. groups only appears if the groups scope is in oidc.scopes and the app's groups claim filter allows it; userinfo fallback cannot fill a claim nobody asked for
Microsoft Entra IDIssuer is https://login.microsoftonline.com/<tenant-id>/v2.0. Groups arrive as Object ID GUIDs, so use GUIDs in managed.policies.match.groups, or use App Roles for readable names. If roles come under roles rather than groups, set oidc.groups_claim: roles
Google WorkspaceIssuer is https://accounts.google.com. The id_token has no groups; configure oidc.google_groups (Admin SDK Directory API via a service account with domain-wide delegation) to use group rules. Otherwise gate with oidc.allowed_email_domains and assign policy with managed.policies.match.email_domain. Google ignores offline_access, so for refresh tokens set oidc.scopes: [openid, profile, email] and oidc.extra_auth_params: { access_type: offline, prompt: consent }

Refresh tokens matter

Refresh tokens let the gateway renew sessions silently, and they are how deprovisioning works: disable someone in the IdP and their next refresh fails, ending the session within ttl_hours. The gateway asks for offline_access by default; if your IdP needs explicit consent for that, allow it on the client.

With no refresh tokens at all, the gateway still works, but people redo the browser login each time the session expires. Raising session.ttl_hours to 8 or 12 stops that happening hourly, at the cost of a disabled user keeping access until the longer TTL runs out.

Deployment

The gateway is one stateless Linux binary coordinating through Postgres. Deploy it like any other stateless internal service that holds a production credential: inside your network, reachable by developers over HTTPS, able to reach your IdP.

Decisions worth making up front:

  • Cost. There is no licence or per-seat fee. It is part of the claude binary, so you pay for inference under your existing agreement plus the compute it runs on.
  • Bypass. The gateway does not stop someone with their own credential calling the provider directly. Closing that is a network decision, for example blocking egress to api.anthropic.com except from the gateway. Doing so also breaks the WebFetch domain safety check from laptops, so set skipWebFetchPreflight: true in the managed policy (see Data usage).
  • More than one gateway. Each is a separate deployment with its own config. The CLI stores trust and credentials per gateway hostname, so teams can use different gateways side by side. One instance serves one OIDC issuer; run more instances for more issuers.
  • Serverless. Cloud Run is fine with min-instances: 1 to avoid cold OIDC discovery. Lambda and Cloud Functions will not work because the gateway is a long-running HTTP server.

Behind a proxy

Every production layout here puts an L7 proxy (an Ingress, Cloud Run's front end, an ALB) in front of plain-HTTP replicas. Set listen.trusted_proxies to the proxy's source ranges so the gateway reads client IPs from X-Forwarded-For; it only honours the header from a trusted peer. Skip this and every request appears to come from the proxy, which squashes per-IP rate limits into one bucket and puts the proxy's IP in every audit event. The AWS and Google Cloud guides have concrete values.

Two more proxy rules:

  • Never redirect the device-authorisation and token endpoints (HTTP to HTTPS, host canonicalisation and so on). Claude Code does not follow redirects on those calls, so sign-in and refresh break.
  • Idle timeout longer than the keepalive. On every upstream except provider: anthropic, the gateway sends an SSE ping after about 15 seconds of silence; on provider: anthropic it passes the API's own pings through. A typical 60-second default is enough. The AWS guide raises the ALB timeout to an hour anyway; gateways older than v2.1.229 sent nothing during quiet periods.

Choosing the address

Claude Code accepts a gateway in one of two ways:

  • Private address. The gateway sits behind an internal load balancer or VPN, and its hostname resolves only to private addresses such as RFC 1918 or CGNAT 100.64.0.0/10. Users can be anywhere. Before you start lists the accepted ranges.
  • Declared block. If your internal network runs on public IPv4 space your organisation owns, list it in the gatewayInternalNetworks managed setting. Both the gateway and the user's machine must sit inside that one block. See gateways on public address space you own.

If no single block covers both, use a private address.

Building the image

Build your own image around the native claude binary:

  1. Download the Linux build for your architecture from a pinned release (see Setup for the URL pattern).
  2. Verify it against the release's GPG-signed manifest.json, also covered in Setup.
  3. Copy it into your build context. Mirror releases to an internal registry if builds cannot reach the release host.

The image also needs:

  • glibc. The glibc build depends only on glibc libraries. On a musl image use the linux-x64-musl or linux-arm64-musl build plus the extra packages described in Setup.
  • A writable state directory. Minimal images have no writable home, so set CLAUDE_CONFIG_DIR to somewhere like /tmp/.claude. Any user works.
  • The command: claude gateway --config /etc/claude/gateway.yaml, with the config mounted read-only and secrets passed as environment variables. It listens on listen.port, default 8080.

A sketch of the Dockerfile I use:

FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y --no-install-recommends ca-certificates \
 && rm -rf /var/lib/apt/lists/*
COPY claude /usr/local/bin/claude
ENV CLAUDE_CONFIG_DIR=/tmp/.claude
USER 10001
EXPOSE 8080
ENTRYPOINT ["claude", "gateway", "--config", "/etc/claude/gateway.yaml"]

Kubernetes

Run it as a Deployment:

  • config from a ConfigMap, secrets from a Secret, referenced in the YAML as ${file:/path/to/secret} or environment variables;
  • TLS terminated at the Ingress, with listen.public_url set to the Ingress hostname;
  • readiness probe on GET /readyz, liveness on GET /healthz.

Prefer workload identity to static keys; the upstreams reference covers each platform. For cross-cloud pairings, such as Bedrock from GKE, give the upstream explicit credentials in its auth block. Deploy on AWS is a complete ECS Fargate or EKS walkthrough.

Cloud Run

  • Leave listen.port at 8080 (Cloud Run's default PORT) or set port: ${PORT}.
  • Set public_url to the origin users actually reach. In production that is usually an internal load balancer's hostname, because /login rejects public addresses and *.run.app resolves to one, so the run.app URL is only good for a smoke test. The exception is a network where *.run.app resolves privately through Private Service Connect and a Cloud DNS private zone.
  • Mount the config as a secret volume.
  • Set min-instances: 1.

Deploy on Google Cloud walks through Cloud Run or GKE end to end.

Pointing laptops at it

Once it is serving, push forceLoginMethod, forceLoginGatewayUrl and parentSettingsBehavior: "merge" through managed settings, via MDM or by writing the per-OS managed-settings.json. Without them, /login shows the normal account picker with no gateway option. As soon as these keys land, Claude Code stops using any leftover API key or claude.ai login on the machine, so ship them alongside your sign-in instructions. The Claude Desktop equivalent, bootstrapUrl, is in client-side managed settings.

Sizing sign-in rate limits for a big launch

Sign-in is rate limited per client IP: by default each address gets 30 sign-in starts and 10 code submissions per 10 minutes. That suits a small team. On a launch morning for thousands, you will hit it if:

  • the gateway cannot see past the load balancer, so everyone shares one address. Fix listen.trusted_proxies first; the gateway logs a warning the first time it ignores an X-Forwarded-For;
  • lots of people share a few NAT or VPN egress addresses. Raise rate_limits (see HTTP tuning).

To size max: developers per shared egress address, times the fraction who sign in within one window_seconds (600 by default), doubled for retries and people signing in to both Claude Code and Claude Desktop. Worked example: 6,000 developers behind 3 egress addresses, signing in evenly over two hours. That is 2,000 per address, about 167 per 10-minute window, doubled to 334; round to 400.

rate_limits:
  device_authorization: { max: 400, window_seconds: 600 }
  device_verify: { max: 400, window_seconds: 600 }

device_verify is what stops anyone guessing another person's sign-in code, so raise it only as far as you need. Codes are 8 characters from a 20-character alphabet and expire in 10 minutes, so guessing remains impractical even at high limits.

With refresh tokens, sessions renew silently and you can lower the limits after launch. Without them, people sign in again every session.ttl_hours, so leave the limits sized for that steady rate.

When the limit trips, Claude Code v2.1.274+ shows The gateway is limiting sign-in attempts right now, a v2.1.274+ gateway shows Too many attempts came from your network address on the verification page, and the gateway logs a sign-in refused line naming the setting to change.

Operations

Logs

Everything goes to stderr in two streams.

Audit events are single-line JSON, one per security-relevant event. Ship stderr to your aggregator. Event names:

config.load, session.mint, session.refresh, device.authorize, device.verify, device.callback, auth.denied, access.denied, access.public_client, inference, managed.serve, desktop_bootstrap.serve, desktop_bootstrap.denied, spend.blocked, admin.denied, admin.limit.upsert, admin.limit.delete.

What some of them carry:

EventNotable fields
session.mint, session.refresh (success)sub, email, client_ip, result
auth.deniedReason, client IP, request path (no identity exists yet)
access.deniedReason and client IP. With reason xff_unparseable, also the unreadable X-Forwarded-For entry; with client_ip_unknown, no IP because the connection had no peer address while an access_control list was set
access.public_clientClient IP of the first request per process from a public address while access_control.allow_cidrs is empty. The request is still served; it is a hint that the gateway may be internet-reachable
inferenceWhich upstream served it and the response status
desktop_bootstrap.deniedReason (not_configured, policy_not_opted_in, no_policy_matched) and identity
admin.deniedClient IP, method, path and reason: invalid_key, bearer_rejected or no_credentials. Key material is never logged

Operational logs are human-readable lines prefixed [gateway] for boot, warnings and upstream errors. CLAUDE_GATEWAY_LOG_LEVEL takes debug, info (default), warn or error. At debug, each sign-in and refresh logs the claim names (never values) from the id_token and any userinfo claims, which is how I debug email_claim and groups_claim without logging personal data. Audit events are emitted regardless of level.

Health checks

  • GET /healthz for liveness.
  • GET /readyz for readiness; it checks the store is reachable. With store.readiness_grace_seconds set, it stays ready for that long after the store stops answering.

Both bypass access_control.allow_cidrs. The OAuth discovery document at /.well-known/oauth-authorization-server only returns 200 once config, OIDC discovery, upstream clients and Postgres migrations have all succeeded, which makes it a good end-to-end boot check.

Concurrency

Each replica sends at most 256 requests upstream at once by default, and a stream holds its slot until it ends. Extra requests queue inside the gateway, so users see slow starts or apparent hangs. On a provider: anthropic upstream, waiting longer than timeouts.upstream_ttfb_ms abandons that upstream and returns 502 if no later upstream takes it.

The startup line containing upstream requests: shows the limit. While over it, a replica logs a warning containing client requests are open, at most once a minute.

To serve more, add replicas or set BUN_CONFIG_MAX_HTTP_REQUESTS (a whole number from 1 to 65535) on the container and restart. A replica fills up at roughly the limit divided by average request duration: with 20-second average streams, 256 slots fill at about 13 requests per second. If you autoscale on CPU, set the target below the CPU level seen when that warning appears, since a queuing replica does not burn CPU.

Warning: Every open or queued request holds memory, including its body. Size container memory for peak concurrency and watch it when you change the limit. A replica killed for running out of memory drops every stream it was carrying.

When dependencies fail

Postgres down:

  • Signed-in developers can keep going in principle: tokens validate locally against the JWT secret and refreshes do not touch the store.
  • New sign-ins fail, because the device flow and its rate-limit counters live in Postgres.
  • Spend limits fail open by default; you can make them fail closed.
  • The catch is readiness. By default /readyz fails as soon as Postgres is unreachable, so every replica goes unready together, and if your platform only routes to ready replicas, all traffic fails. /healthz keeps passing.

To ride out a short outage such as a failover, set store.readiness_grace_seconds above your failover time, say 300. With spend limits failing open, traffic through a still-ready replica is unmetered meanwhile, so keep the value tight. With enforcement.fail_closed_on_error: true, signed-in inference gets 429 spend limit unavailable until Postgres returns, even on ready replicas. The grace setting needs v2.1.282+ on the gateway; older gateways refuse to boot with the key present, so upgrade all replicas first. Pointing readiness at /healthz instead also keeps replicas in rotation, but then a replica whose database connection never recovers keeps passing too.

IdP down: existing sessions last until ttl_hours; new logins fail; refreshes get a try-again answer and succeed once the IdP is back. If your IdP has frequent maintenance windows, lengthen ttl_hours.

Rotating the JWT secret

Rotate in stages so nobody is logged out:

  1. Generate a new secret and put it first in the session.jwt_secret array.
  2. Roll the deployment. New tokens use the new secret; old ones still verify.
  3. After ttl_hours plus a margin, remove the old secret and roll again.

There is no per-session revocation, because tokens validate locally. Replacing the secret outright (old one removed) is the only way to kick everyone out at once. To offboard one person, deprovision them in the IdP; their session ends within ttl_hours.

Postgres

Use genuine PostgreSQL, self-hosted or managed, at the minimum version in Before you start. Postgres-protocol-compatible distributed databases are not supported. store.postgres_url takes a single host, so for multi-node databases point it at the managed endpoint, load balancer or virtual IP, and set a readiness grace period longer than a failover.

Tables, all created by boot-time migrations (plus _migrations):

TableHoldsKept for
kvDevice grants (10-minute TTL) and rate-limit countersPer-row TTL
spendPeriod-to-date spend per principal, in centsadmin.spend_retention_months, default 13
spend_limitsConfigured capsUntil deleted through the API
admin_auditAdmin API change trailadmin.audit_retention_days, default 365
principal_emailsLast-seen email, display name and IdP groups per principal. Personal dataadmin.identity_retention_days since last activity, default 90

A 30-second loop expires kv rows and an hourly sweep enforces the other windows. Without spend limits, only kv is written. Because migrations run at every boot, the database role needs rights to create and alter tables; give the gateway its own database or schema to keep that grant narrow.

Back it up if you use spend limits: losing it loses spend tracking and caps, not just sessions. To erase one leaver immediately, DELETE FROM principal_emails WHERE principal = '<sub>' removes the only table holding their email, name and groups. spend and admin_audit only hold the pseudonymous OIDC sub.

Upgrades

Replicas are stateless, so a rolling restart loses nothing. The new binary migrates the schema on boot, and concurrent replicas serialise on a Postgres advisory lock so each migration runs once.

On SIGTERM (rolling restart, scale-in), a v2.1.274+ gateway stops taking new connections and lets in-flight requests and streams finish, for up to 25 seconds. SIGINT drains the same way; a second signal exits immediately. To allow long generations more time on Kubernetes or ECS, raise both together:

  • CLAUDE_GATEWAY_DRAIN_TIMEOUT_MS on the container, as a plain positive integer such as 180000. Anything else (like 180s) is ignored and the 25-second default stays.
  • The orchestrator grace period: terminationGracePeriodSeconds on Kubernetes, stopTimeout on ECS. Both default to 30 seconds. Keep it at least 5 seconds above the drain window, and on Kubernetes add any preStop hook time, since the grace clock starts before the hook.

Platform ceilings: ECS on Fargate caps stopTimeout at 120 seconds, and Cloud Run kills an instance 10 seconds after SIGTERM whatever you set. When the drain window expires with requests still open, the gateway logs a warning containing drain window over after, the number of cut requests, and both settings to raise.

Rollback is safe for the database: migrations are append-only and an older binary ignores rows it does not know. But the older binary validates the YAML against its own schema, so remove any keys introduced by the newer release before rolling back.

Patching: you pin the version in your own image, so fixes, security fixes included, reach you only when you bump the pin and redeploy. Put the gateway on the same patch cadence as other services holding production credentials.

Security

Data flow

DataPathDoes the gateway send it to Anthropic?
Prompts and completionsCLI, gateway, your upstreamOnly if the Anthropic API is a configured upstream
Telemetry (OTLP metrics, plus opt-in logs and traces)CLI, gateway, your collectorNever
Identity (email, groups, sub)IdP, gateway, CLI, which stamps it on OTLP exports. With forward_user_identity, also sent as headers to your own proxyNever
Managed settingsYour YAML, CLINever
Audit logGateway stderr, your aggregatorNever

Threat model

The gateway is inside your perimeter but does not trust individual laptops:

  • Developers hold short-lived JWTs, never upstream keys. The CLI-to-gateway leg uses the RFC 8628 device grant, and the gateway's code exchange with the IdP uses PKCE by default, so an intercepted authorisation code is useless.
  • The device verification page requires same-origin POSTs and applies a per-IP rate limit (RFC 8628 section 5.1).
  • Requests to your IdP, OTLP collectors and provider: anthropic upstreams go through an SSRF guard: it resolves DNS, blocks link-local, cloud metadata and (by default) loopback addresses, and pins the connection to the resolved IP. RFC 1918 ranges are allowed on purpose, since IdPs and collectors often live there. For other providers, a base_url naming one of those addresses or a metadata hostname is refused at config load, and the provider SDK then connects without the DNS check. With proxy-only egress, the check moves to your forward proxy, whose allowlist must refuse those destinations.
  • CLAUDE_GATEWAY_ALLOW_LOOPBACK=1 relaxes the loopback block for every operator URL and skips the boot warning about metadata reachability. Use it only for genuinely local things like a dev IdP or a sidecar collector, and prefer giving those an internal address.

If you add egress controls, keep the metadata server reachable whenever the gateway uses instance-metadata credentials such as workload identity.

Out of scope, because they are your infrastructure:

  • A compromised gateway host. It holds the upstream credential and pushes managed settings to every developer, so treat its configuration as you would your MDM. The CLI's approval dialog for shell-capable settings (see Server-managed settings) limits silent changes but is no substitute for host security.
  • A malicious IdP. It signs the id_tokens the gateway trusts and can assert any identity.

Sign-in code guessing

The user_code typed on /device is 8 characters from a 20-character alphabet, about 2.56 x 10^10 combinations, valid for 10 minutes, behind per-IP limits on the device-grant endpoints. Those limits apply only to sign-in, never to inference.

Questionnaire answers

  • Data residency. The gateway sends nothing to Anthropic unless the Anthropic API is an upstream, in which case your existing agreement covers inference. Telemetry, audit, identity and settings go only where you configure.
  • The host process is the Claude Code CLI. claude gateway follows the same third-party rules as Bedrock and Agent Platform deployments and sends nothing to Anthropic. Before v2.1.227 it sent startup telemetry (switch off with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 in the container) and a single body-less, credential-less HEAD /api/hello at boot to https://api.anthropic.com or ANTHROPIC_BASE_URL, unless a proxy variable or mTLS client certificate was set. The response was ignored, so blocking it was harmless.
  • Client analytics. The CLI disables its usage analytics and error reporting while signed in to a gateway. Before the first sign-in it still sends startup events, even on machines forced to gateway login; deliver DISABLE_TELEMETRY in the same client-side managed settings to stop them.
  • Error reporting is off whenever model requests go anywhere other than Anthropic's first-party API.
  • Client machines still send WebFetch hostname checks and version checks to Anthropic unless CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 and skipWebFetchPreflight: true are set. Plugin marketplaces have their own switches (below).
  • Survey ratings are not sent to Anthropic while signed in to a gateway, and a Yes on the transcript-share prompt writes a local file under ~/.claude/feedback-bundles/.
  • Updates. Update checks are separate from gateway traffic. Pin versions through your own distribution and set DISABLE_UPDATES if laptops must never fetch releases. DISABLE_AUTOUPDATER only stops background updates; claude update still works.
  • TLS. Serve public_url over HTTPS in production, either with listen.tls or from a TLS-terminating ingress, with listen.public_url set either way. The gateway will not refuse plain HTTP, so this is on you. The IdP must use HTTPS, Postgres supports ?sslmode=require, and set Strict-Transport-Security at the ingress.
  • Vulnerabilities: report as described in Security.

Plugin marketplace traffic

Laptops fetch plugin marketplaces directly, not through the gateway; Network configuration lists the hosts. On a developer's first interactive terminal session, Claude Code registers the official claude-plugins-official marketplace, downloading from downloads.claude.ai and falling back to cloning from github.com.

CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 does not stop that first registration. These managed settings do:

  • a strictKnownMarketplaces allowlist that omits it, or a blockedMarketplaces entry naming it (see Plugins for your organisation);
  • CLAUDE_CODE_DISABLE_OFFICIAL_MARKETPLACE_AUTOINSTALL set to "1" in the managed env block.

Because the first registration can happen before gateway sign-in, put your choice in client-side managed settings as well as in the policy's cli block.

Troubleshooting

When raising a problem with Anthropic support, on the Claude Code GitHub issues, or with your account team, include:

  • Gateway problems: stderr for the window, gateway.yaml with secrets redacted, the version (shown on the landing page at / and in the x-cc-gateway-version header on /managed/settings), and recent changes.
  • Login problems: a debug file from claude --debug-file ./claude-debug.txt plus the gateway audit log for the same window.
  • Inference problems: the model, the configured upstreams, and the audit entry for the request.

Redact before posting publicly: stderr contains audit events with identities, and the debug file contains hook and MCP output.

Sign-in and connection errors

You seeCauseFix
/login shows the normal account picker, not Cloud gatewayforceLoginMethod or forceLoginGatewayUrl missing on that machineDeploy the managed settings; see Point machines at the gateway
Not signed in to the Cloud gateway, followed by a prompt to run /loginManaged settings require gateway sign-in and the session has none; a leftover claude.ai login does not countRun /login and complete gateway sign-in
Gateway login is configured in managed settings, but this Claude Code build does not include Cloud gateway support.Client predates gateway supportUpdate Claude Code
Administrator policy requires a Cloud gateway sign-in on this machine at startupANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN set, an apiKeyHelper configured, or an old Console key savedUnset the variable, remove the helper, or run claude auth logout; then /login. Sessions selecting a cloud provider with CLAUDE_CODE_USE_* start without sign-in
Claude Code may not be enabled for your organization after a 403 on managed settingsSomething answered /managed/settings with 403; the gateway's own route never does. It is the access_control IP checks or a proxy or WAF in frontLook for access.denied in the audit log; fix access_control or the front end
The gateway is limiting sign-in attempts right now (older: Request failed with status code 429)Per-IP sign-in limit reached; audit events show result: rate_limited from few IPsFix listen.trusted_proxies, then raise rate_limits (sizing)
Gateway hosts must be on your organization's private network; <host> resolves to the public (or unrecognized) address <ip>At least one resolved address is public. Dual-stack names are a common trap, including AWS internal dual-stack load balancers returning public-range AAAA recordsMake the name resolve only to private addresses, or serve an internal-only name; or declare your owned block
Gateway login would go through proxy <proxy>, which is not on a private networkAn HTTPS_PROXY/HTTP_PROXY with a public hostname applies to the gatewayAdd the gateway to NO_PROXY (the message names the entry) or use a privately resolved proxy
Claude Code only signs in to <host> from inside its declared network <block> ...Machine reached a declared-block gateway from outside the block (VPN pool, container or WSL2 NAT, foreign network)Sign in from the host OS on your network; widen the block (up to /8) if the address is also yours. A second overlapping entry is refused
Every address for gateway host <host> must be inside its declared network <block> ...A record outside the block, such as a second site or an IPv6 recordPublish only in-block records, or use a separate internal name
<host> is on the declared network <block>, which Claude Code checks over a direct connection, not through an HTTP proxyA proxy applies to a declared-block gatewayAdd the NO_PROXY entry named
Message starting gatewayInternalNetworks in managed settingsThe value breaks a validation rule; all new gateway logins on the machine are refused until fixedCorrect the entry named
Could not resolve the configured HTTP proxy / Could not resolve gateway host <host>Machine is off the corporate networkConnect to the network or VPN, or fix the URL
Claude Desktop cannot fetch its bootstrap configuration/user/bootstrap returned 404: no matching policy, or it lacks a desktop key (desktop_bootstrap.denied in the audit log)Add a desktop block (even desktop: {}) to the matching policy or the match: {} base layer; see Claude Desktop overlay

Boot and database errors

You seeCauseFix
Config validation error naming store.postgres_urlNo Postgres configuredSet it. Locally: docker run --rm -p 5432:5432 -e POSTGRES_HOST_AUTH_METHOD=trust postgres
store.postgres_url in <path> is not a URL the gateway can read (before v2.1.290: Invalid URL or URI error)Multiple hosts, or a password with unencoded /, ?, # or %Use one host and move the password to store.password
requires the native binaryRunning under NodeUse a standalone install (Setup)
OIDC discovery error after config.loadIssuer unreachable or its TLS chain untrustedCheck reachability; set ca_cert_pem; behind a forward proxy set oidc.use_proxy: true (see outbound proxies), or proxy-only egress (v2.1.277+) if DNS or IP CONNECT is also blocked
Postgres permission error at bootRole lacks DDL rightsGrant CREATE on the gateway's schema
could not connect to Postgres at boot, attempt 1 of 3Database not yet reachableHarmless if boot then completes. It tries three times, two seconds apart. If it exits, check the URL and network; if attempts time out, raise store.connect_timeout_seconds

IdP and session errors

You seeCauseFix
/oauth/callback says "Sign-in could not be completed"Email domain rejected, id_token invalid, or email_verified is false (always rejected)Check allowed_email_domains and the IdP's verified email; set oidc.email_claim for a different claim name
token exchange failed ... id_token missing email claimIdP omits email (only rejected when allowed_email_domains is set)Emit email in the id_token (Okta custom server claim, Entra optional claim, PingFederate OIDC policy), or set oidc.userinfo_fallback: true
refresh failed ... invalid_token (...) (at userinfo_no_id_token, ...) and Cloud gateway session expired every TTLIdP returned no id_token on refresh and then rejected the refreshed access token at userinfoSet oidc.scope_on_refresh: true (v2.1.260+). On PingFederate, enable Return ID Token On Refresh Grant instead. As a stopgap, raise session.ttl_hours
Unknown or unsupported scopeIdP rejects a scopeSet oidc.scopes to exactly what it accepts, including openid. Default is openid profile email offline_access
No silent renewal after overriding scopesoffline_access droppedAdd it back if supported
"This request came from another site and was blocked"Cross-site form POST blocked as CSRFOpen the verification link directly
Chrome blocks Approve with a form-action CSP errorIdP redirects via a second hostAdd each origin in the chain to oidc.form_action_origins
Callback fails with CSP error (Chrome) or "this sign-in link has expired" (Safari)IdP used response_mode=form_postMake the IdP honour the response_mode=query the gateway requests
Works locally, fails behind an ALBpublic_url still names the inner http:// originSet listen.public_url to the external https:// origin and register its callback

TLS, load and header errors

You seeCauseFix
Trust prompt keeps reappearingCertificate differs per replica or requestTerminate TLS once with a stable certificate
"Could not verify the gateway's TLS certificate" or SELF_SIGNED_CERT_IN_CHAINPrivate CA not trusted on the laptopThe native binary (and Node 22.15+) reads the OS store, controlled by CLAUDE_CODE_CERT_STORE; otherwise set NODE_EXTRA_CA_CERTS. See Network configuration
Cloud gateway sign-in was not completed with a certificate mismatchFirst request after sign-in saw a different certificate than the pinned oneServe one certificate per hostname and /login again; the trust prompt reappears with a changed-certificate warning. The message shows the hostname and the first 16 characters of both fingerprints
The gateway's TLS certificate changed during sign-inDifferent certificate mid-sign-in: mixed replicas, interception, or rotationServe one certificate, then restart sign-in
Every Bedrock request 502s with Could not load credentials from any providersOn EC2, IMDSv2 hop limit 1 blocks metadata from the container; boot passes because credentials load lazilyRaise the hop limit to 2 (aws ec2 modify-instance-metadata-options --http-put-response-hop-limit 2), ideally on a dedicated instance, or use ECS task roles
Slow starts, hangs, or 502 all upstreams failed at peak with a healthy upstreamReplica concurrency limit reached; log shows client requests are openAdd replicas or raise the limit (concurrency)
Every request after sign-in fails with 431Headers too large because the token lists many IdP groupsSee below

If Claude Code reports couldn't load your organization's managed settings after sign-in, it names the reason, restarts in place and resumes; where it cannot restart (a background session, say), it ends the session but keeps the sign-in.

431 after sign-in

The session token in each request's Authorization header lists the developer's IdP groups. The gateway answers 431 when headers exceed 256 KiB, or limits.max_request_header_bytes if set (before v2.1.284 the limit was 16 KiB), and logs nothing. In order:

  1. Gateway older than v2.1.284: upgrade.
  2. limits.max_request_header_bytes set: raise or remove it.
  3. Otherwise: have the IdP emit fewer groups, keeping any named in oidc.allowed_groups, admin.admin_groups, match.groups in managed.policies, and rbac_group spend caps, since those decide access, policy and caps.