Deploy Claude apps gateway on Google Cloud
A worked Google Cloud build of Claude apps gateway on Cloud Run or GKE, with Cloud SQL, Secret Manager and a service account calling Agent Platform.
This walkthrough runs Claude apps gateway on Google Cloud, with Google Cloud's Agent Platform (formerly Vertex AI) as the upstream and either Cloud Run or GKE for compute. It is a working example of customer-managed infrastructure, not a supported production blueprint, so use it to understand the moving parts and then fit it to your estate. Platform-neutral requirements are in Deploying and running Claude apps gateway.
Google Workspace is the example IdP. Any OIDC provider works and only the oidc block changes; see the identity provider section.
What gets built
| Component | Purpose |
|---|---|
| Cloud Run service or GKE Deployment | Runs the gateway |
| Artifact Registry repository | Holds the image |
| Cloud SQL for PostgreSQL, private IP only | The gateway's store |
| Secret Manager | gateway.yaml, JWT key, OIDC client secret, Postgres URL |
Service account with roles/aiplatform.user | Attached directly on Cloud Run, via Workload Identity on GKE |
| HTTPS front end | An internal Application Load Balancer in front of Cloud Run (you create it; this page only configures the gateway for it), or a GKE internal Ingress of class gce-internal |
Prerequisites
- A project with billing enabled and rights to create the above.
gcloudauthenticated withgcloud auth login, and Docker.- For GKE:
kubectland a cluster on the VPC created below. - Access to the Claude models you need in Model Garden, in a region that publishes them.
- A Google Workspace OAuth 2.0 web client with redirect URI
https://<gateway-host>/oauth/callback. - A TLS hostname for the gateway, normally an internal DNS name for the load balancer.
export PROJECT_ID=acme-ai-platform
export REGION=us-east5 # a region where your Claude models are published
gcloud config set project "$PROJECT_ID"
Step 1: enable APIs
gcloud services enable aiplatform.googleapis.com artifactregistry.googleapis.com \
sqladmin.googleapis.com secretmanager.googleapis.com iamcredentials.googleapis.com \
iam.googleapis.com compute.googleapis.com servicenetworking.googleapis.com \
run.googleapis.com container.googleapis.com
compute and servicenetworking are for the private-IP Cloud SQL path; run is Cloud Run only and container is GKE only, so drop whichever you are not using.
Step 2: service account
The gateway runs as its own service account allowed to call Agent Platform. It talks to Cloud SQL over the VPC with a password user, so it needs no Cloud SQL IAM role.
gcloud iam service-accounts create cc-gateway --display-name="Claude apps gateway"
SA="cc-gateway@${PROJECT_ID}.iam.gserviceaccount.com"
gcloud projects add-iam-policy-binding "$PROJECT_ID" \
--member="serviceAccount:${SA}" --role="roles/aiplatform.user" --condition=None
Then enable the Claude models for the project in Model Garden. Each model publishes to specific regions, so check every model card against $REGION.
Step 3: build and push the image
Build to the image requirements with the linux-x64 glibc binary:
gcloud artifacts repositories create cc-gateway --repository-format=docker --location="$REGION"
gcloud auth configure-docker "${REGION}-docker.pkg.dev" --quiet
IMAGE="${REGION}-docker.pkg.dev/${PROJECT_ID}/cc-gateway/gateway:2.1.290"
docker build --platform=linux/amd64 --provenance=false -t "$IMAGE" .
docker push "$IMAGE"
Cloud Run only runs linux/amd64, and --provenance=false stops buildx producing an OCI image index, which Cloud Run rejects.
Step 4: Cloud SQL on a private VPC
Private Services Access gives the instance a private IP and no public one, which also keeps you compliant where constraints/sql.restrictPublicIp is enforced.
VPC=ccgw-vpc
gcloud compute networks create "$VPC" --subnet-mode=custom
gcloud compute networks subnets create ccgw-subnet --network="$VPC" \
--region="$REGION" --range=10.60.0.0/24
# Private Services Access, once per VPC
gcloud compute addresses create "google-managed-services-${VPC}" \
--global --purpose=VPC_PEERING --prefix-length=16 --network="$VPC"
gcloud services vpc-peerings connect --service=servicenetworking.googleapis.com \
--ranges="google-managed-services-${VPC}" --network="$VPC"
gcloud sql instances create ccgw-db --database-version=POSTGRES_16 --tier=db-g1-small \
--region="$REGION" --network="projects/${PROJECT_ID}/global/networks/${VPC}" --no-assign-ip
gcloud sql databases create claude_gateway --instance=ccgw-db
DB_PASSWORD="$(openssl rand -hex 24)"
gcloud sql users create gateway --instance=ccgw-db --password="$DB_PASSWORD"
DB_IP="$(gcloud sql instances describe ccgw-db --format='value(ipAddresses[0].ipAddress)')"
GATEWAY_POSTGRES_URL="postgres://gateway:${DB_PASSWORD}@${DB_IP}:5432/claude_gateway?sslmode=require"
Whichever runtime you choose has to sit on, or be routed into, this VPC.
Step 5: gateway.yaml
The Agent Platform upstream uses auth: {}, meaning Application Default Credentials from the runtime's service account. Two listen fields describe the front end:
public_url: the externalhttps://origin, required for any non-loopback bind. The IdPredirect_uriand the discovery document are built only from this, never fromX-Forwarded-*.trusted_proxies: the front end's source ranges.X-Forwarded-Foris honoured only from these peers, so rate limits and audit events show developer IPs rather than the load balancer's.
| Front end | trusted_proxies |
|---|---|
| Cloud Run reached directly, no load balancer | [169.254.0.0/16] |
| Internal Application Load Balancer in front of Cloud Run | 169.254.0.0/16 plus your proxy-only subnet CIDR |
GKE internal Ingress, class gce-internal | Your proxy-only subnet CIDR |
An external GKE Ingress (class gce) is deliberately absent: it gets a public forwarding-rule address, and /login rejects public addresses (see Before you start).
This example assumes an internal load balancer in front of Cloud Run:
listen:
host: 0.0.0.0
port: 8080
public_url: https://claude.gw.acme-health.internal
trusted_proxies: [169.254.0.0/16, 10.129.0.0/23]
oidc:
issuer: https://accounts.google.com
client_id: 123456789012-abcdefg.apps.googleusercontent.com
client_secret: ${OIDC_CLIENT_SECRET} # GKE: ${file:/secrets/oidc-client-secret}
allowed_email_domains: [acme-health.co.uk]
scopes: [openid, profile, email] # Google ignores offline_access
extra_auth_params: { access_type: offline, prompt: consent }
session:
jwt_secret: ${GATEWAY_JWT_SECRET} # GKE: ${file:/secrets/jwt-secret}
store:
postgres_url: ${GATEWAY_POSTGRES_URL} # GKE: ${file:/secrets/postgres-url}
# readiness_grace_seconds: 300 # stay ready through a Cloud SQL failover
upstreams:
- provider: vertex
region: us-east5 # same as $REGION
project_id: acme-ai-platform
auth: {}
The extra_auth_params line is what gets you refresh tokens from Google, and therefore silent session renewal.
Note: Google id_tokens have no
groupsclaim. To drivemanaged.policiesby group with Workspace as the IdP, configureoidc.google_groups, which looks groups up through the Admin SDK Directory API using a service account with domain-wide delegation. Otherwise match onemail_domain.
Step 6: secrets
Create four secrets and grant roles/secretmanager.secretAccessor on them to the gateway's service account:
| Secret | Value |
|---|---|
gateway-jwt-secret | openssl rand -base64 32 |
gateway-oidc-client-secret | From the OAuth client in the Google Cloud console |
gateway-postgres-url | $GATEWAY_POSTGRES_URL |
gateway-config | The whole gateway.yaml |
For example:
openssl rand -base64 32 | gcloud secrets create gateway-jwt-secret --data-file=-
printf '%s' "$GATEWAY_POSTGRES_URL" | gcloud secrets create gateway-postgres-url --data-file=-
gcloud secrets create gateway-config --data-file=gateway.yaml
Delivery differs by track:
- GKE: everything mounts as files through the Secret Manager CSI driver, and
gateway.yamluses${file:/secrets/...}. - Cloud Run: it cannot mount several secrets into one directory, so
gateway.yamlmounts as a file and the other three arrive as environment variables, referenced as${GATEWAY_JWT_SECRET},${OIDC_CLIENT_SECRET}and${GATEWAY_POSTGRES_URL}.
Step 7: deploy
Option A: Cloud Run
Production deployment behind an internal load balancer:
gcloud run deploy cc-gateway \
--image="$IMAGE" --region="$REGION" \
--service-account="cc-gateway@${PROJECT_ID}.iam.gserviceaccount.com" \
--min-instances=1 --max-instances=8 --timeout=3600 \
--ingress=internal \
--network="$VPC" --subnet=ccgw-subnet --vpc-egress=private-ranges-only \
--set-secrets=/etc/claude/gateway.yaml=gateway-config:latest,GATEWAY_JWT_SECRET=gateway-jwt-secret:latest,OIDC_CLIENT_SECRET=gateway-oidc-client-secret:latest,GATEWAY_POSTGRES_URL=gateway-postgres-url:latest \
--no-invoker-iam-check
Why each choice:
- Direct VPC egress (
--network,--subnet,--vpc-egress=private-ranges-only) reaches Cloud SQL's private IP. Public traffic to Agent Platform andaccounts.google.comgoes straight out, so no Cloud NAT is needed. --max-instances=8. Each instance holds up tostore.max_connectionsPostgres connections, five by default. Keep max instances times that figure under your Cloud SQL tier's connection limit; 8 suitsdb-g1-small.--timeout=3600. Cloud Run's default request timeout is 300 seconds, which would cut long streams.--min-instances=1avoids cold OIDC discovery.- Invoker check open. Gateway clients carry no Google token, so Cloud Run's invoker check must admit unauthenticated requests; the gateway's own OIDC sign-in, with
allowed_email_domains, does the authentication.--no-invoker-iam-checkdisables it without anallUsersbinding and works under Domain Restricted Sharing. If your organisation forbids that flag,--allow-unauthenticatedgrantsallUserstherun.invokerrole instead. --ingress=internalis an independent layer from the invoker check; keep it to restrict the service to your network.
Giving it a private hostname. The *.run.app URL normally resolves to a public address, which /login refuses. Cloud Run creates neither of the two fixes for you:
- Internal Application Load Balancer (what this page's YAML assumes). Put one in front with an internal DNS name and certificate, and set
listen.public_urlto it.--ingress=internalalready admits internal load balancers.internal-and-cloud-load-balancingwould also admit external load balancers, whose public addresses/loginrejects, so you do not need it. - Internal ingress, no load balancer. Leave
public_urlas the*.run.appURL. That only works if your network team already runs a Private Service Connect endpoint for Google APIs, a Cloud DNS private zone resolving*.run.appto it, and on-premises routing to that endpoint.
Google's private networking guide for Cloud Run describes both. Until the private hostname is in place, check the container booted from its Cloud Run logs rather than trying to sign in.
Update the OAuth client's redirect URI to <public_url>/oauth/callback before the first sign-in, and redeploy whenever public_url changes; the gateway ignores X-Forwarded-Host and X-Forwarded-Proto and only honours X-Forwarded-For when trusted_proxies is set.
Option B: GKE
The cluster must be on $VPC itself. Peering another VPC to it will not work, because Cloud SQL private IP is already a peered network and peering is not transitive. For a new cluster, pass --network="$VPC" --subnetwork=ccgw-subnet to gcloud container clusters create.
Enable Workload Identity and bind the Google service account to a Kubernetes one:
gcloud container clusters update ai-tools --region="$REGION" \
--workload-pool="${PROJECT_ID}.svc.id.goog"
# Standard clusters: existing node pools need GKE_METADATA (Autopilot has it already)
gcloud container node-pools update default-pool --cluster=ai-tools \
--region="$REGION" --workload-metadata=GKE_METADATA
kubectl create namespace claude-gateway
kubectl create serviceaccount gateway -n claude-gateway
gcloud iam service-accounts add-iam-policy-binding "$SA" \
--role roles/iam.workloadIdentityUser \
--member "serviceAccount:${PROJECT_ID}.svc.id.goog[claude-gateway/gateway]"
kubectl annotate serviceaccount gateway -n claude-gateway iam.gke.io/gcp-service-account="$SA"
Then deploy a Deployment, Service and gce-internal Ingress as in the Kubernetes section, with serviceAccountName: gateway, secrets mounted at /secrets by the CSI driver, and readiness on GET /readyz.
The backend service behind a GKE Ingress times out after 30 seconds by default, which kills streaming responses. Attach a BackendConfig to the Service:
apiVersion: cloud.google.com/v1
kind: BackendConfig
metadata:
name: gateway-backend
namespace: claude-gateway
spec:
timeoutSec: 3600
healthCheck:
requestPath: /readyz
port: 8080
---
apiVersion: v1
kind: Service
metadata:
name: gateway
namespace: claude-gateway
annotations:
cloud.google.com/backend-config: '{"default": "gateway-backend"}'
spec:
selector: { app: claude-gateway }
ports: [{ port: 80, targetPort: 8080 }]
Do not add an egress NetworkPolicy blocking 169.254.169.254 on a Workload Identity cluster: pods need the metadata server for credentials, and the gateway's own SSRF guard (see the threat model) is the protection there. For the same reason, the boot warning that the metadata endpoint is reachable is expected under Workload Identity and can be ignored.
Step 8: point laptops at it
Push the full managed settings snippet (forceLoginMethod, forceLoginGatewayUrl and parentSettingsBehavior: "merge") to every device via MDM, as in point machines at the gateway. There is no manual gateway option in the login picker.
Reference assets
The Claude Code GitHub repository has an examples/gateway/gcp bundle that automates the Cloud Run track (the config and image assets also work for GKE):
setup.sh, an idempotentgcloudscript from enabling APIs to first deploy;terraform/, the same deployment as code for a greenfield project: a targeted apply to create the Artifact Registry repository, then build and push, then the full apply;gateway.yaml.exampleand aDockerfilefor a distroless runtime image.
The bundle also defaults Cloud Run to internal ingress and leaves public_url as the *.run.app URL, and it does not create a load balancer. One difference from this page: it opens the invoker layer with an allUsers run.invoker grant rather than --no-invoker-iam-check. Either works; pick whichever your organisation policies allow. As with this page, it is an example to review and adapt.
Google Cloud troubleshooting
The general table is in the deployment guide.
| Symptom | Cause | Fix |
|---|---|---|
Cloud Run answers 403 Forbidden before the container sees anything | Invoker IAM check still on | Deploy with --no-invoker-iam-check, or --allow-unauthenticated |
--no-invoker-iam-check fails with invoker_iam_disabled is not currently available | constraints/run.managed.requireInvokerIam is enforced | Use --allow-unauthenticated. If Domain Restricted Sharing (constraints/iam.allowedPolicyMemberDomains) blocks that as well, use the GKE track, which needs no allUsers binding |
Container manifest type … must support amd64/linux on deploy | Built on a non-amd64 host or buildx produced an OCI index | Rebuild with --platform=linux/amd64 --provenance=false |
| Postgres connection timeout at boot on Cloud Run | Service not on the VPC, or Cloud SQL lacks a private IP there | Use --network and --subnet for Direct VPC egress; create Cloud SQL with --no-assign-ip on the same VPC |
Agent Platform returns 403 PERMISSION_DENIED | Runtime is not using the gateway's service account, or the model is not enabled in Model Garden | Set --service-account (Cloud Run) or bind Workload Identity (GKE); enable each model for the region |
| Streams cut off after a fixed time | Front-end timeout: 30 seconds on a GKE Ingress backend, 300 seconds on Cloud Run | BackendConfig timeoutSec on GKE; --timeout=3600 on Cloud Run |