Claude apps gateway
Run Claude Code through Bedrock, Google Cloud, Foundry or the Anthropic API behind a self-hosted gateway with SSO sign-in, per-group policy and OTLP telemetry.
The Claude apps gateway is a server you host yourself that sits between your developers' Claude Code clients and whichever model provider you pay for. Developers sign in with their corporate SSO account, and the gateway holds the real upstream credential. Nobody on the team ever sees an API key or a cloud access key.
The server ships inside the normal claude binary. The same executable you run on a laptop becomes the gateway when you start it with claude gateway --config gateway.yaml.
This page explains when the gateway is worth running, walks through a first deployment end to end, and then covers how developer machines connect and what the gateway does and does not support. The full YAML reference lives in Claude apps gateway configuration, and caps per developer are in spend limits.
Note: The gateway is aimed at organisations that need inference to run through their own cloud account, typically for data residency or procurement reasons. If that is not a constraint for you, a Claude Enterprise plan gives you SCIM provisioning, Claude Code on the web and mobile, and less to operate. The feature availability page compares the deployment routes.
What the gateway gives you
The general case for putting a gateway in front of Claude Code is covered in gateways. What this particular one adds is that it is built and released alongside Claude Code itself, so it already knows every header and body field the client sends. You do not maintain a passthrough allowlist.
Once it is running you get five things:
| Capability | What it means in practice |
|---|---|
| Credential custody | The upstream key or cloud role lives only on your servers. Developers hold short-lived bearer tokens minted after SSO. Disable someone in the IdP and their access lapses within the session lifetime (one hour by default). |
| Group-based access | IdP groups map to model allowlists and full managed settings documents. Model access is checked server-side, and the settings are applied by the CLI at the managed tier, so developers cannot override locked keys. |
| Settings delivery | The gateway serves managed settings to signed-in clients itself, replacing server-managed settings from the claude.ai admin console. |
| Telemetry relay | Each configured destination receives OTLP metrics stamped with user identity, model, tokens and latency. Logs and traces are opt-in per destination. See monitoring usage. |
| Upstream routing | Clients always speak the Anthropic Messages API. The gateway translates for Amazon Bedrock, Claude Platform on AWS, Google Cloud's Agent Platform, Microsoft Foundry or the Anthropic API, and fails over between them. You can move regions or providers without touching a single laptop. |
The gateway's own data plane talks to Anthropic only if you configure the Anthropic API as an upstream. Telemetry, audit logs, managed settings and identity information go only where you send them. The CLI process itself can still make some non-inference calls; the deployment guide covers how to close those off.
When to use something else
If you already operate an LLM proxy or API gateway that works for you, carry on using it and follow other LLM gateways. The gateway protocol page lists exactly what Claude Code expects from any gateway.
A running Claude apps gateway also describes its own client-facing protocol at GET /protocol. That document covers sign-in, inference, settings delivery, model discovery and telemetry endpoints. Breaking changes are announced in advance, but backwards compatibility is not promised forever.
Before you start
Gather these before writing any YAML:
- Claude Code v2.1.195 or newer on the gateway host and on every developer machine. The
claude gatewaysubcommand and the gateway login flow arrived in that release. The Claude Platform on AWS upstream needs v2.1.198 on the server. Runclaude updateif in doubt. - An OpenID Connect identity provider. Okta, Microsoft Entra ID, Google Workspace, Keycloak, Dex, PingFederate and any other compliant IdP work. SAML and LDAP do not.
- PostgreSQL 11 or later. It stores device sign-in state and rate-limit counters, and with spend limits enabled it also stores spend, audit and identity tables you should back up. The smallest managed tier is fine. Use
?sslmode=requirefor TLS. Versions 11 to 13 need v2.1.290 on the server, and since they are out of upstream support you should prefer something newer. - A model upstream credential: an AWS role, Claude Platform on AWS key, Google Cloud service account, Foundry resource or Anthropic API key. You can list several.
- HTTPS. Laptops and the browser used for sign-in must reach the gateway over
https://. Terminate TLS on the gateway withlisten.tlsor at an ingress, and setlisten.public_urleither way. Plainhttp://is accepted at/loginonly forlocalhost,127.0.0.1or::1. - A private address. At
/login, Claude Code insists the gateway hostname resolves only to private ranges: RFC 1918, link-local, CGNAT100.64.0.0/10, IPv6 ULAfc00::/7or loopback. This matters because a trusted gateway can push settings that execute commands on developer machines. - Linux for the server. Only the native Linux build runs the gateway in production. macOS is fine for local experiments. Windows is not a supported server platform.
Warning: If laptops send HTTPS through a corporate proxy, sign-in also requires the proxy host to resolve privately. If it does not, add the gateway host to
NO_PROXYso the CLI connects directly.
A first deployment
The route below gets you from nothing to a developer signed in, using Docker Compose and an Amazon Bedrock upstream. Swapping in another provider only changes the upstreams block, which the configuration reference documents.
1. Register an OAuth client
Pick the gateway hostname first, because the IdP redirect URI must match it exactly. Create an OIDC web application with the redirect URI https://<your-gateway-host>/oauth/callback. Keep the client ID and secret to hand. Per-IdP walkthroughs are in the deployment guide.
2. Create the database
Provision PostgreSQL. The gateway applies its own schema migrations every time it boots, so the role you give it needs permission to create and alter tables.
3. Write gateway.yaml
Five sections are mandatory and everything else has a default. Secrets go in as ${ENV_VAR} references so the file can sit in version control. Here is a starting point for a company whose IdP is Entra ID and whose AWS account lives in Ireland:
listen:
host: 0.0.0.0
port: 8080
public_url: https://claude.tools.northwind.internal
oidc:
issuer: https://login.microsoftonline.com/7d1e0c55-1111-4b2a-9c3d-0f0f0f0f0f0f/v2.0
client_id: 4c2b9a10-2222-4e7f-8a1b-123456789abc
client_secret: ${ENTRA_CLIENT_SECRET}
allowed_email_domains: [northwind.co.uk]
userinfo_fallback: true
session:
jwt_secret: ${GW_SIGNING_SECRET}
ttl_hours: 1
store:
postgres_url: ${GW_DATABASE_URL}
upstreams:
- provider: bedrock
region: eu-west-1
auth: {}
auto_include_builtin_models: true
auth: {} tells the gateway to use the AWS default credential chain, which picks up IRSA, an ECS task role, an instance profile, environment variables or ~/.aws. auto_include_builtin_models: true exposes the built-in catalogue of Claude models with their Bedrock IDs already mapped. For a non-US region like this one you will usually add a models: block with the right regional IDs; the configuration reference shows how.
Note: The AWS principal needs
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStreamon both the inference-profile ARNs and the underlying foundation-model ARNs, and someone in the account must have submitted Anthropic's one-time use case form from the Bedrock model catalogue. Prefer an attached role over static keys.
Generate the signing secret with openssl rand -base64 32.
4. Run it
Wrap the claude binary in a container image that meets the image requirements in the deployment guide, then run it beside Postgres. A Compose file for a local trial:
services:
claude-gw:
image: ghcr.io/northwind/claude-gw:2.1.290
ports: ["8080:8080"]
volumes:
- ./gateway.yaml:/etc/claude/gateway.yaml:ro
command: ["claude", "gateway", "--config", "/etc/claude/gateway.yaml"]
environment:
ENTRA_CLIENT_SECRET: ${ENTRA_CLIENT_SECRET}
GW_SIGNING_SECRET: ${GW_SIGNING_SECRET}
GW_DATABASE_URL: postgres://claudegw:localonly@db/claudegw
# Local trial only: borrow your own short-lived AWS session
AWS_ACCESS_KEY_ID: ${AWS_ACCESS_KEY_ID}
AWS_SECRET_ACCESS_KEY: ${AWS_SECRET_ACCESS_KEY}
AWS_SESSION_TOKEN: ${AWS_SESSION_TOKEN}
depends_on:
db:
condition: service_healthy
db:
image: postgres:17
environment:
POSTGRES_USER: claudegw
POSTGRES_PASSWORD: localonly
POSTGRES_DB: claudegw
healthcheck:
test: ["CMD-SHELL", "pg_isready -U claudegw"]
interval: 5s
In production, drop any static AWS credentials and let the workload role supply them.
Boot is fail-closed. The gateway reads the config, connects to Postgres, applies migrations, runs OIDC discovery and builds the upstream clients. If any of those steps fails it exits rather than serving traffic half-configured. Note that Bedrock and Google Cloud credentials are only resolved on the first real request, so a clean boot does not prove inference works.
Watch stderr while it starts. Operational lines look like [gateway] <timestamp> <level> <message>, audit events are single-line JSON with an evt field, and a fresh database prints one migration N applied line per migration. The line you want is claude gateway listening on http://0.0.0.0:8080. You will also see a warning that access_control.allow_cidrs is empty; that is expected until you add an allow list.
If it exits before the listening line, the last line of stderr names the cause. The usual suspects are an unreachable database, a role without DDL rights, a broken OIDC discovery document or a schema error in the YAML (reported with the field path).
If you already have a TLS-terminating ingress, skip Compose: run claude gateway --config gateway.yaml directly, bind listen to a loopback or cluster-internal address, and set public_url to the ingress origin.
5. Check the sign-in path
Three quick checks, in order, tell you where a problem is. For a local trial without an ingress, use http://localhost:8080, set public_url to the same value, and add http://localhost:8080/oauth/callback as a second redirect URI on the OAuth client. On Windows PowerShell type curl.exe, because plain curl is an alias for Invoke-WebRequest.
# 1. Discovery document: proves boot completed
curl -s https://claude.tools.northwind.internal/.well-known/oauth-authorization-server | jq .issuer
# 2. Device authorisation: proves Postgres is reachable and writable
curl -s -X POST https://claude.tools.northwind.internal/oauth/device_authorization | jq .verification_uri_complete
The second call returns a device_code, a short user_code, a verification_uri_complete, a 600 second expiry and a 5 second polling interval. For the third check, open verification_uri_complete in a browser, confirm the code, sign in at your IdP and make sure you land back on a confirmation page.
| First check that fails | Where to look |
|---|---|
| Discovery | Boot never finished. Read stderr. |
| Device authorisation | Database connection string or grants. |
| Browser never reaches the IdP | Redirect URI does not exactly match https://<gateway>/oauth/callback. |
| IdP bounces back with an error | The gateway audit log records every rejected sign-in with a reason, for example email domain not allowed. |
6. Sign a developer in
On a developer laptop, put forceLoginMethod: "gateway" and forceLoginGatewayUrl (your public_url) in the machine's managed settings file, run /login, press Enter on the Cloud gateway screen and finish the browser step. The next section explains how to roll that out to everyone.
Connecting developer machines
Developers need no claude.ai account, no subscription and no API key. Everything is driven by managed settings you push from your MDM, so there is nothing for them to configure.
Point machines at the gateway
Push these keys in the per-OS managed settings file:
{
"forceLoginMethod": "gateway",
"forceLoginGatewayUrl": "https://claude.tools.northwind.internal",
"parentSettingsBehavior": "merge"
}
The two login keys open /login straight on the Cloud gateway screen with your URL filled in. parentSettingsBehavior: "merge" lets Claude Desktop hand the gateway's egress allowlist to the Claude Code sessions it launches (more on that below).
Some things to know:
- A developer cannot set this up themselves. The login picker has no gateway option, and
forceLoginGatewayUrlis ignored in user settings.forceLoginMethodon its own leaves them at a "Contact your IT administrator" message. - The login keys belong in the file on the device, not in the gateway's
managed.policies[].cliblock, because that block only reaches clients that are already connected. - Developers who choose a cloud provider via a variable such as
CLAUDE_CODE_USE_BEDROCKare not forced through the gateway sign-in. To close that route, add"allowedProviders": ["gateway"]as described in client-side managed settings. - A machine that has the file but no gateway sign-in shows one of the messages under errors.
Certificate pinning on first connect
The first time a client talks to the gateway it records the SHA-256 fingerprint of the TLS leaf certificate and pins it to that hostname. The pin is rechecked at sign-in, on silent session refreshes and when fetching managed settings. Ordinary inference uses standard TLS validation. Requests sent via an HTTPS proxy skip the pin check, which is another reason to put the gateway in NO_PROXY.
The /login prompt shows the first 16 characters of the fingerprint in lowercase hex without colons. Publish the full value so developers can compare. This prints it in the same format:
openssl x509 -in gateway-leaf.pem -noout -fingerprint -sha256 \
| cut -d= -f2 | tr -d ':' | tr '[:upper:]' '[:lower:]'
Every certificate rotation re-triggers the trust prompt for everyone, so schedule rotations and republish the fingerprint. Approval of security-sensitive managed settings is remembered against the pinned certificate too, so those approval dialogs also reappear after a rotation.
A gateway may include an email field in its token response so the developer can confirm which account they used before the credential is saved (Claude Code v2.1.275 or newer). The gateway inside the claude binary does not send it, so its sign-ins skip that confirmation. /status shows the account after a confirmed sign-in.
What happens after sign-in
- The model picker shows only the models in the developer's
availableModelslist. - Managed settings apply at startup and refresh hourly.
- Telemetry flows to your collector via the gateway.
- The session refreshes silently before
ttl_hoursruns out. If the refresh fails because the user was deprovisioned, Claude Code asks them to log in again.
Gateways on public address space you own
Some networks are numbered from public IPv4 blocks the organisation owns. For those, list the blocks in the gatewayInternalNetworks managed setting (Claude Code v2.1.268 or newer). /login will then accept a gateway inside a listed block when the developer's machine also connects from inside that same block.
{
"gatewayInternalNetworks": ["198.20.0.0/16"]
}
The rules Claude Code enforces on the list:
- IPv4 only, written as the first address of the block with a prefix between
/8and/32. - At most four blocks, none overlapping each other.
- No overlap with private space (
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,127.0.0.0/8,169.254.0.0/16,100.64.0.0/10), because those already work. - No overlap with ranges that are never a real network:
198.18.0.0/15,192.0.0.0/24, the documentation ranges192.0.2.0/24,198.51.100.0/24and203.0.113.0/24, plus0.0.0.0/8,192.88.99.0/24and multicast224.0.0.0/4. Blocks within240.0.0.0/4are allowed.
Entries from managed-settings.json and its managed-settings.d/ drop-ins are combined before validation. Only managed file, MDM profile or registry policy sources are honoured; user, project and server-managed settings are ignored. If the value is invalid, Claude Code refuses every new gateway sign-in on that machine, including private ones, while existing sessions keep working. Test on one machine first.
With a valid list, three checks apply to a gateway inside a declared block: every address its hostname resolves to must sit in that one block, the developer's own source address must be in the same block (so NAT, containers, WSL2 and off-block VPN pools fail), and the connection must be direct rather than through HTTPS_PROXY. When all pass, the trust prompt names both addresses and the block.
Warning: This setting does not make an internet-facing gateway safe. Keep the gateway unreachable from outside with firewall rules, set
access_control.allow_cidrsto the same blocks, and setlisten.trusted_proxieswhen a load balancer sits in front so the allow list sees real client addresses. Never declare a block shared with other tenants, such as a cloud provider's public range.
Policy for Claude Desktop sessions
Claude Desktop runs its Cowork and Code tabs (and Chat, if enabled) on embedded Claude Code sessions that send model requests through the gateway. Desktop builds their policy from what the gateway serves at /user/bootstrap: the model allowlist, disabled tools and egress allowlist from the matched policy, plus the policy's desktop overlay. Other cli keys such as hooks, env and scoped permission rules like Bash(npm *) only reach clients that sign in with /login.
Desktop applies the model list and disabled tools itself, but the egress allowlist reaches embedded sessions only as parent settings (as WebFetch domain rules and sandbox network rules). Claude Code ignores parent settings on any machine with an admin-deployed managed source unless the selected source sets parentSettingsBehavior: "merge". Without it, embedded sessions silently run without the egress restriction. A marketplace allowlist (strictKnownMarketplaces) sent by Desktop 2.16120.0 or newer is delivered the same way.
So:
- Desktop-only machines need the opt-in. It is already in the snippet above.
- Machines signing in via
/logindo not, since each session fetches policy directly. policyHelperfleets cannot use it, because Claude Code reads managed settings only from the helper's output and never merges parent settings.
Claude Code reads parentSettingsBehavior only from the selected managed source. A macOS managed-preferences plist or Windows HKLM policy outranks the file, and the gateway's own remote settings outrank both, so mirror the whole snippet into any higher source you use and into the policy's cli block. To confirm which client-side source wins on a Desktop-only machine, call the Agent SDK's resolveSettings() and read policyOrigin on the managed entry: it reports plist, hklm or file.
Locking down parent settings
Once merge is on, any process that launches Claude Code can supply parent settings: Desktop, an Agent SDK app or an IDE extension. Claude Code filters them to an allowlist of restrictive keys, but some of those keys can widen access, so pair the opt-in with the five allowManaged*Only locks and your own allowlists:
{
"forceLoginMethod": "gateway",
"forceLoginGatewayUrl": "https://claude.tools.northwind.internal",
"parentSettingsBehavior": "merge",
"allowManagedPermissionRulesOnly": true,
"allowManagedMcpServersOnly": true,
"allowManagedHooksOnly": true,
"allowedMcpServers": [{ "serverUrl": "https://mcp.northwind.internal/*" }],
"sandbox": {
"network": {
"allowManagedDomainsOnly": true,
"allowedDomains": ["dev.azure.com", "*.pypi.org", "files.pythonhosted.org"]
},
"filesystem": {
"allowManagedReadPathsOnly": true,
"denyRead": ["~/"],
"allowRead": ["~/src"]
}
}
}
Points to keep in mind:
- Each lock ignores the developer's own entries for that setting, so always ship the matching allowlist. An empty managed domain list with the network lock blocks all sandboxed egress. The MCP lock with no
allowedMcpServersanywhere loads every server thatdeniedMcpServersdoes not block.allowReadonly re-opens paths inside adenyReadregion. - Parent-supplied
sandbox.credentialsentries arrive stripped:denyentries keep only path or name and mode, filemaskentries become sentinel-only whole-file masks with emptyinjectHosts,envVarsmask entries are dropped, andawsPairs/sigv4are forwarded restriction-only. - Some keys still pass the filter with every lock on:
forceLoginOrgUUID,allowedMcpServers,availableModels,allowedProviders,strictKnownMarketplaces,blockedMarketplaces(additive) andstrictPluginOnlyCustomization(no lock blocks it). For the list-type keys, set your own value in the winning admin source and the parent's is ignored. - Under the default first-wins behaviour most locks need to be in the winning source. With the
managedSourcesBehaviormerge opt-in described in managed settings, the strictest value from any source applies.
Connecting Claude Desktop itself
Desktop uses its own MDM key: set bootstrapUrl in its managed configuration to <public_url>/user/bootstrap, and opt the user's policy in with a desktop key on the gateway (server v2.1.203 or newer). Desktop then runs the same browser SSO and fetches its configuration from your gateway rather than from Anthropic. The CLI and Desktop sign in separately; they do not share a session. To turn on the Chat tab, set chatTabEnabled: true in Desktop's managed configuration or in the policy's desktop block (server v2.1.227 or newer). See the Claude Desktop overlay.
CI and headless machines
There is no service-token flow. Sign-in is always the browser device flow, so an unattended CI job cannot authenticate through the gateway; point those jobs at your provider directly, as in GitHub Actions with cloud providers.
Once a person has signed in on a machine, every Claude Code process there uses that session, including claude -p and Agent SDK runs, and the gateway policy applies to each. Because the device flow separates the polling CLI from the approving browser, a headless dev box works fine: run /login over SSH and open the link on your laptop.
What is enforced on developers
For sessions signed in through /login:
- Models. Requests for an ungranted model get a 400, and
/modelonly lists the allowed ones. This includes the default model a session starts on, so see start sessions on an allowed model. - Telemetry. OTLP/HTTP exports go to the gateway, not a locally set
OTEL_EXPORTER_OTLP_ENDPOINT, unless a policy names your collector directly. Signals with no configured destination are accepted and discarded. Desktop's embedded sessions send to the configured endpoint and attach the gateway token only when that endpoint is the gateway. - Credentials. The gateway token is the only credential used. Anthropic profiles and earlier claude.ai logins are ignored while signed in. A configured
ANTHROPIC_API_KEY,ANTHROPIC_AUTH_TOKENorapiKeyHelpertriggers the error described in errors. - Managed settings. Locked keys cannot be overridden. Changes land on the hourly poll, apart from the ones that only apply at next launch.
- Gateway unreachable at startup. Signed-in sessions exit after about 10 seconds rather than running without policy.
- Deprovisioning. A disabled user's session ends within
ttl_hourswhen the next refresh fails. - Sign-out.
/logoutdeletes the local credential. If the gateway advertises arevocation_endpointon the same origin, the client also revokes the tokens there (v2.1.275 or newer). The built-in gateway advertises none, so to force sessions out server-side, rotate the JWT secret as described in the deployment guide.
What the organisation can see
Telemetry carries each developer's identity, token counts, model and latency to your collector. The gateway never logs or stores prompt or completion content. Logs and traces, which can include commands and file paths, are only collected if you enable them for a destination.
Feature support through the gateway
The gateway forwards every anthropic-beta value the CLI sends. Bedrock ignores the header, so for that upstream the values are moved into the body's anthropic_beta field.
| Feature | Status | Notes |
|---|---|---|
| Inference to Bedrock, Claude Platform on AWS, Agent Platform, Foundry, Anthropic | Yes | Per-upstream model translation and failover. Bedrock uses bedrock-runtime. The Mantle upstream needs server v2.1.283; Claude Platform on AWS needs v2.1.198. |
| Model access and managed settings by IdP group | Yes | Models enforced server-side, settings applied at the managed tier. |
| Claude Desktop | Yes, opt-in | Needs a desktop key in the policy and server v2.1.203. |
| OTLP/HTTP telemetry fan-out | Yes | Protobuf and JSON encodings. |
| Any OIDC IdP | Yes | One issuer per gateway. |
| Per-user and per-group spend limits | Yes | See spend limits. |
| Standard prompt caching | Yes | cache_control breakpoints reach every upstream. See prompt caching. |
| Auto mode | Yes | Follows the third-party provider model rules in permission modes. Before v2.1.207 it needed CLAUDE_CODE_ENABLE_AUTO_MODE=1. |
| Server-side web search | No | The CLI cannot tell which upstream serves a request, so WebSearch is disabled. |
| Remote Control | No | Shows an error naming the gateway. |
/design-sync and /design-login | No | They need claude.ai, which gateway sessions never contact. |
Features needing feature-flag fetches, such as /import and claude import | No | Flag fetching is skipped. See environment variables. |
| 1-hour cache TTL | No | Gateway sessions use the 5-minute TTL. |
| Global cache scope, token-efficient tools and similar first-party optimisations | No | Not enabled on gateway sessions. |
| OTLP over gRPC | No | HTTP only. |
| SAML, LDAP | No | Put an OIDC bridge in front if you must. |
| Multiple OIDC issuers | No | Run one gateway per issuer. |
| Windows server, Helm chart, admin UI | No | Linux only; it runs as a plain stateless Deployment; configuration is the YAML file. |
Where to go next
- Grow
gateway.yamlwith group policies, failover and telemetry using the configuration reference. - Move from Compose to Kubernetes or Cloud Run and harden it with the deployment guide.
- Put a ceiling on each developer with spend limits.
- Follow a full worked build on AWS or Google Cloud.