Skip to content

Deploy Claude apps gateway on Google Cloud

A worked Google Cloud build of Claude apps gateway on Cloud Run or GKE, with Cloud SQL, Secret Manager and a service account calling Agent Platform.

This walkthrough runs Claude apps gateway on Google Cloud, with Google Cloud's Agent Platform (formerly Vertex AI) as the upstream and either Cloud Run or GKE for compute. It is a working example of customer-managed infrastructure, not a supported production blueprint, so use it to understand the moving parts and then fit it to your estate. Platform-neutral requirements are in Deploying and running Claude apps gateway.

Google Workspace is the example IdP. Any OIDC provider works and only the oidc block changes; see the identity provider section.

What gets built

ComponentPurpose
Cloud Run service or GKE DeploymentRuns the gateway
Artifact Registry repositoryHolds the image
Cloud SQL for PostgreSQL, private IP onlyThe gateway's store
Secret Managergateway.yaml, JWT key, OIDC client secret, Postgres URL
Service account with roles/aiplatform.userAttached directly on Cloud Run, via Workload Identity on GKE
HTTPS front endAn internal Application Load Balancer in front of Cloud Run (you create it; this page only configures the gateway for it), or a GKE internal Ingress of class gce-internal

Prerequisites

  • A project with billing enabled and rights to create the above.
  • gcloud authenticated with gcloud auth login, and Docker.
  • For GKE: kubectl and a cluster on the VPC created below.
  • Access to the Claude models you need in Model Garden, in a region that publishes them.
  • A Google Workspace OAuth 2.0 web client with redirect URI https://<gateway-host>/oauth/callback.
  • A TLS hostname for the gateway, normally an internal DNS name for the load balancer.
export PROJECT_ID=acme-ai-platform
export REGION=us-east5        # a region where your Claude models are published
gcloud config set project "$PROJECT_ID"

Step 1: enable APIs

gcloud services enable aiplatform.googleapis.com artifactregistry.googleapis.com \
  sqladmin.googleapis.com secretmanager.googleapis.com iamcredentials.googleapis.com \
  iam.googleapis.com compute.googleapis.com servicenetworking.googleapis.com \
  run.googleapis.com container.googleapis.com

compute and servicenetworking are for the private-IP Cloud SQL path; run is Cloud Run only and container is GKE only, so drop whichever you are not using.

Step 2: service account

The gateway runs as its own service account allowed to call Agent Platform. It talks to Cloud SQL over the VPC with a password user, so it needs no Cloud SQL IAM role.

gcloud iam service-accounts create cc-gateway --display-name="Claude apps gateway"
SA="cc-gateway@${PROJECT_ID}.iam.gserviceaccount.com"
gcloud projects add-iam-policy-binding "$PROJECT_ID" \
  --member="serviceAccount:${SA}" --role="roles/aiplatform.user" --condition=None

Then enable the Claude models for the project in Model Garden. Each model publishes to specific regions, so check every model card against $REGION.

Step 3: build and push the image

Build to the image requirements with the linux-x64 glibc binary:

gcloud artifacts repositories create cc-gateway --repository-format=docker --location="$REGION"
gcloud auth configure-docker "${REGION}-docker.pkg.dev" --quiet

IMAGE="${REGION}-docker.pkg.dev/${PROJECT_ID}/cc-gateway/gateway:2.1.290"
docker build --platform=linux/amd64 --provenance=false -t "$IMAGE" .
docker push "$IMAGE"

Cloud Run only runs linux/amd64, and --provenance=false stops buildx producing an OCI image index, which Cloud Run rejects.

Step 4: Cloud SQL on a private VPC

Private Services Access gives the instance a private IP and no public one, which also keeps you compliant where constraints/sql.restrictPublicIp is enforced.

VPC=ccgw-vpc
gcloud compute networks create "$VPC" --subnet-mode=custom
gcloud compute networks subnets create ccgw-subnet --network="$VPC" \
  --region="$REGION" --range=10.60.0.0/24

# Private Services Access, once per VPC
gcloud compute addresses create "google-managed-services-${VPC}" \
  --global --purpose=VPC_PEERING --prefix-length=16 --network="$VPC"
gcloud services vpc-peerings connect --service=servicenetworking.googleapis.com \
  --ranges="google-managed-services-${VPC}" --network="$VPC"

gcloud sql instances create ccgw-db --database-version=POSTGRES_16 --tier=db-g1-small \
  --region="$REGION" --network="projects/${PROJECT_ID}/global/networks/${VPC}" --no-assign-ip
gcloud sql databases create claude_gateway --instance=ccgw-db
DB_PASSWORD="$(openssl rand -hex 24)"
gcloud sql users create gateway --instance=ccgw-db --password="$DB_PASSWORD"

DB_IP="$(gcloud sql instances describe ccgw-db --format='value(ipAddresses[0].ipAddress)')"
GATEWAY_POSTGRES_URL="postgres://gateway:${DB_PASSWORD}@${DB_IP}:5432/claude_gateway?sslmode=require"

Whichever runtime you choose has to sit on, or be routed into, this VPC.

Step 5: gateway.yaml

The Agent Platform upstream uses auth: {}, meaning Application Default Credentials from the runtime's service account. Two listen fields describe the front end:

  • public_url: the external https:// origin, required for any non-loopback bind. The IdP redirect_uri and the discovery document are built only from this, never from X-Forwarded-*.
  • trusted_proxies: the front end's source ranges. X-Forwarded-For is honoured only from these peers, so rate limits and audit events show developer IPs rather than the load balancer's.
Front endtrusted_proxies
Cloud Run reached directly, no load balancer[169.254.0.0/16]
Internal Application Load Balancer in front of Cloud Run169.254.0.0/16 plus your proxy-only subnet CIDR
GKE internal Ingress, class gce-internalYour proxy-only subnet CIDR

An external GKE Ingress (class gce) is deliberately absent: it gets a public forwarding-rule address, and /login rejects public addresses (see Before you start).

This example assumes an internal load balancer in front of Cloud Run:

listen:
  host: 0.0.0.0
  port: 8080
  public_url: https://claude.gw.acme-health.internal
  trusted_proxies: [169.254.0.0/16, 10.129.0.0/23]

oidc:
  issuer: https://accounts.google.com
  client_id: 123456789012-abcdefg.apps.googleusercontent.com
  client_secret: ${OIDC_CLIENT_SECRET}        # GKE: ${file:/secrets/oidc-client-secret}
  allowed_email_domains: [acme-health.co.uk]
  scopes: [openid, profile, email]            # Google ignores offline_access
  extra_auth_params: { access_type: offline, prompt: consent }

session:
  jwt_secret: ${GATEWAY_JWT_SECRET}           # GKE: ${file:/secrets/jwt-secret}

store:
  postgres_url: ${GATEWAY_POSTGRES_URL}       # GKE: ${file:/secrets/postgres-url}
  # readiness_grace_seconds: 300              # stay ready through a Cloud SQL failover

upstreams:
  - provider: vertex
    region: us-east5                          # same as $REGION
    project_id: acme-ai-platform
    auth: {}

The extra_auth_params line is what gets you refresh tokens from Google, and therefore silent session renewal.

Note: Google id_tokens have no groups claim. To drive managed.policies by group with Workspace as the IdP, configure oidc.google_groups, which looks groups up through the Admin SDK Directory API using a service account with domain-wide delegation. Otherwise match on email_domain.

Step 6: secrets

Create four secrets and grant roles/secretmanager.secretAccessor on them to the gateway's service account:

SecretValue
gateway-jwt-secretopenssl rand -base64 32
gateway-oidc-client-secretFrom the OAuth client in the Google Cloud console
gateway-postgres-url$GATEWAY_POSTGRES_URL
gateway-configThe whole gateway.yaml

For example:

openssl rand -base64 32 | gcloud secrets create gateway-jwt-secret --data-file=-
printf '%s' "$GATEWAY_POSTGRES_URL" | gcloud secrets create gateway-postgres-url --data-file=-
gcloud secrets create gateway-config --data-file=gateway.yaml

Delivery differs by track:

  • GKE: everything mounts as files through the Secret Manager CSI driver, and gateway.yaml uses ${file:/secrets/...}.
  • Cloud Run: it cannot mount several secrets into one directory, so gateway.yaml mounts as a file and the other three arrive as environment variables, referenced as ${GATEWAY_JWT_SECRET}, ${OIDC_CLIENT_SECRET} and ${GATEWAY_POSTGRES_URL}.

Step 7: deploy

Option A: Cloud Run

Production deployment behind an internal load balancer:

gcloud run deploy cc-gateway \
  --image="$IMAGE" --region="$REGION" \
  --service-account="cc-gateway@${PROJECT_ID}.iam.gserviceaccount.com" \
  --min-instances=1 --max-instances=8 --timeout=3600 \
  --ingress=internal \
  --network="$VPC" --subnet=ccgw-subnet --vpc-egress=private-ranges-only \
  --set-secrets=/etc/claude/gateway.yaml=gateway-config:latest,GATEWAY_JWT_SECRET=gateway-jwt-secret:latest,OIDC_CLIENT_SECRET=gateway-oidc-client-secret:latest,GATEWAY_POSTGRES_URL=gateway-postgres-url:latest \
  --no-invoker-iam-check

Why each choice:

  • Direct VPC egress (--network, --subnet, --vpc-egress=private-ranges-only) reaches Cloud SQL's private IP. Public traffic to Agent Platform and accounts.google.com goes straight out, so no Cloud NAT is needed.
  • --max-instances=8. Each instance holds up to store.max_connections Postgres connections, five by default. Keep max instances times that figure under your Cloud SQL tier's connection limit; 8 suits db-g1-small.
  • --timeout=3600. Cloud Run's default request timeout is 300 seconds, which would cut long streams.
  • --min-instances=1 avoids cold OIDC discovery.
  • Invoker check open. Gateway clients carry no Google token, so Cloud Run's invoker check must admit unauthenticated requests; the gateway's own OIDC sign-in, with allowed_email_domains, does the authentication. --no-invoker-iam-check disables it without an allUsers binding and works under Domain Restricted Sharing. If your organisation forbids that flag, --allow-unauthenticated grants allUsers the run.invoker role instead.
  • --ingress=internal is an independent layer from the invoker check; keep it to restrict the service to your network.

Giving it a private hostname. The *.run.app URL normally resolves to a public address, which /login refuses. Cloud Run creates neither of the two fixes for you:

  • Internal Application Load Balancer (what this page's YAML assumes). Put one in front with an internal DNS name and certificate, and set listen.public_url to it. --ingress=internal already admits internal load balancers. internal-and-cloud-load-balancing would also admit external load balancers, whose public addresses /login rejects, so you do not need it.
  • Internal ingress, no load balancer. Leave public_url as the *.run.app URL. That only works if your network team already runs a Private Service Connect endpoint for Google APIs, a Cloud DNS private zone resolving *.run.app to it, and on-premises routing to that endpoint.

Google's private networking guide for Cloud Run describes both. Until the private hostname is in place, check the container booted from its Cloud Run logs rather than trying to sign in.

Update the OAuth client's redirect URI to <public_url>/oauth/callback before the first sign-in, and redeploy whenever public_url changes; the gateway ignores X-Forwarded-Host and X-Forwarded-Proto and only honours X-Forwarded-For when trusted_proxies is set.

Option B: GKE

The cluster must be on $VPC itself. Peering another VPC to it will not work, because Cloud SQL private IP is already a peered network and peering is not transitive. For a new cluster, pass --network="$VPC" --subnetwork=ccgw-subnet to gcloud container clusters create.

Enable Workload Identity and bind the Google service account to a Kubernetes one:

gcloud container clusters update ai-tools --region="$REGION" \
  --workload-pool="${PROJECT_ID}.svc.id.goog"
# Standard clusters: existing node pools need GKE_METADATA (Autopilot has it already)
gcloud container node-pools update default-pool --cluster=ai-tools \
  --region="$REGION" --workload-metadata=GKE_METADATA

kubectl create namespace claude-gateway
kubectl create serviceaccount gateway -n claude-gateway

gcloud iam service-accounts add-iam-policy-binding "$SA" \
  --role roles/iam.workloadIdentityUser \
  --member "serviceAccount:${PROJECT_ID}.svc.id.goog[claude-gateway/gateway]"
kubectl annotate serviceaccount gateway -n claude-gateway iam.gke.io/gcp-service-account="$SA"

Then deploy a Deployment, Service and gce-internal Ingress as in the Kubernetes section, with serviceAccountName: gateway, secrets mounted at /secrets by the CSI driver, and readiness on GET /readyz.

The backend service behind a GKE Ingress times out after 30 seconds by default, which kills streaming responses. Attach a BackendConfig to the Service:

apiVersion: cloud.google.com/v1
kind: BackendConfig
metadata:
  name: gateway-backend
  namespace: claude-gateway
spec:
  timeoutSec: 3600
  healthCheck:
    requestPath: /readyz
    port: 8080
---
apiVersion: v1
kind: Service
metadata:
  name: gateway
  namespace: claude-gateway
  annotations:
    cloud.google.com/backend-config: '{"default": "gateway-backend"}'
spec:
  selector: { app: claude-gateway }
  ports: [{ port: 80, targetPort: 8080 }]

Do not add an egress NetworkPolicy blocking 169.254.169.254 on a Workload Identity cluster: pods need the metadata server for credentials, and the gateway's own SSRF guard (see the threat model) is the protection there. For the same reason, the boot warning that the metadata endpoint is reachable is expected under Workload Identity and can be ignored.

Step 8: point laptops at it

Push the full managed settings snippet (forceLoginMethod, forceLoginGatewayUrl and parentSettingsBehavior: "merge") to every device via MDM, as in point machines at the gateway. There is no manual gateway option in the login picker.

Reference assets

The Claude Code GitHub repository has an examples/gateway/gcp bundle that automates the Cloud Run track (the config and image assets also work for GKE):

  • setup.sh, an idempotent gcloud script from enabling APIs to first deploy;
  • terraform/, the same deployment as code for a greenfield project: a targeted apply to create the Artifact Registry repository, then build and push, then the full apply;
  • gateway.yaml.example and a Dockerfile for a distroless runtime image.

The bundle also defaults Cloud Run to internal ingress and leaves public_url as the *.run.app URL, and it does not create a load balancer. One difference from this page: it opens the invoker layer with an allUsers run.invoker grant rather than --no-invoker-iam-check. Either works; pick whichever your organisation policies allow. As with this page, it is an example to review and adapt.

Google Cloud troubleshooting

The general table is in the deployment guide.

SymptomCauseFix
Cloud Run answers 403 Forbidden before the container sees anythingInvoker IAM check still onDeploy with --no-invoker-iam-check, or --allow-unauthenticated
--no-invoker-iam-check fails with invoker_iam_disabled is not currently availableconstraints/run.managed.requireInvokerIam is enforcedUse --allow-unauthenticated. If Domain Restricted Sharing (constraints/iam.allowedPolicyMemberDomains) blocks that as well, use the GKE track, which needs no allUsers binding
Container manifest type … must support amd64/linux on deployBuilt on a non-amd64 host or buildx produced an OCI indexRebuild with --platform=linux/amd64 --provenance=false
Postgres connection timeout at boot on Cloud RunService not on the VPC, or Cloud SQL lacks a private IP thereUse --network and --subnet for Direct VPC egress; create Cloud SQL with --no-assign-ip on the same VPC
Agent Platform returns 403 PERMISSION_DENIEDRuntime is not using the gateway's service account, or the model is not enabled in Model GardenSet --service-account (Cloud Run) or bind Workload Identity (GKE); enable each model for the region
Streams cut off after a fixed timeFront-end timeout: 30 seconds on a GKE Ingress backend, 300 seconds on Cloud RunBackendConfig timeoutSec on GKE; --timeout=3600 on Cloud Run