Deploy Claude apps gateway on AWS
A worked AWS build of Claude apps gateway: ECS Fargate or EKS, RDS for PostgreSQL, Secrets Manager, an internal ALB and IAM-role access to Bedrock.
This is one way to stand Claude apps gateway up on AWS, with Amazon Bedrock as the model upstream. Treat it as a working example of customer-managed infrastructure rather than a supported production blueprint: follow it to see how the parts connect, then adapt it. The platform-neutral rules live in Deploying and running Claude apps gateway.
Compute is either Amazon ECS on AWS Fargate or Amazon EKS. Okta is the example IdP, but any OIDC provider works; the identity provider section covers the others.
Note: Bedrock is not the only Claude upstream on AWS. Claude Platform on AWS (the Anthropic-operated Claude API with AWS authentication and AWS Marketplace billing) can replace Bedrock or sit alongside it. Its upstream entry, credentials and IAM permissions differ; see
provider: anthropicAws. Everything else here applies unchanged.
What gets built
Developers sign in through Okta to a private HTTPS endpoint. Their sessions reach Claude on Bedrock using the gateway's IAM role, so no model credentials ever reach a laptop.
| Component | Role |
|---|---|
| ECS Fargate service or EKS Deployment | Runs the gateway container |
| Amazon ECR repository | Holds the image |
| Amazon RDS for PostgreSQL, private subnets, not publicly accessible | The gateway's store |
| AWS Secrets Manager | JWT signing key, OIDC client secret, Postgres URL |
IAM role with bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:CountTokens | ECS task role, or bound through IRSA on EKS |
| Internal Application Load Balancer | HTTPS front end |
Prerequisites
The walkthrough creates the gateway's own resources on top of infrastructure you already have:
- an AWS account allowed to create the resources above;
- AWS CLI v2, authenticated, and Docker;
- a VPC with at least two private subnets in different Availability Zones and outbound access through a NAT gateway (the internal ALB needs two AZs, and the gateway needs to reach Bedrock and the IdP);
- an Okta OIDC web app with redirect URI
https://<gateway-host>/oauth/callback; - a TLS hostname for the gateway, typically in a Route 53 private hosted zone aliased to the ALB, with an ACM certificate for it, imported or issued by AWS Private CA.
Shell variables
Every command reads four values. Choose a US region where Bedrock serves the models you want: the gateway's built-in model catalogue resolves to us.anthropic.* inference profiles, and the IAM policy below grants those ARNs. Outside the US, add a models block with your geography's inference profile IDs and change the policy's ARN prefix to match.
Find two private subnets if you need to:
aws ec2 describe-subnets --filters "Name=vpc-id,Values=vpc-0a1b2c3d4e5f60718" \
--query 'Subnets[].[SubnetId,AvailabilityZone,CidrBlock]' --output text
Then export:
export AWS_REGION=us-west-2
export ACCOUNT_ID="$(aws sts get-caller-identity --query Account --output text)"
export VPC_ID=vpc-0a1b2c3d4e5f60718
export PRIVATE_SUBNETS="subnet-0aa11bb22cc33dd44 subnet-0ee55ff66aa77bb88"
Step 1: security groups
Three groups form a chain: corporate network to ALB on 443, ALB to gateway on 8080, gateway to Postgres on 5432. Nothing else gets in.
ALB_SG="$(aws ec2 create-security-group --group-name ccgw-alb \
--description "Claude gateway load balancer" --vpc-id "$VPC_ID" --query GroupId --output text)"
GW_SG="$(aws ec2 create-security-group --group-name ccgw-app \
--description "Claude gateway tasks" --vpc-id "$VPC_ID" --query GroupId --output text)"
DB_SG="$(aws ec2 create-security-group --group-name ccgw-db \
--description "Claude gateway database" --vpc-id "$VPC_ID" --query GroupId --output text)"
aws ec2 authorize-security-group-ingress --group-id "$ALB_SG" --protocol tcp --port 443 --cidr 10.40.0.0/14
aws ec2 authorize-security-group-ingress --group-id "$GW_SG" --protocol tcp --port 8080 --source-group "$ALB_SG"
aws ec2 authorize-security-group-ingress --group-id "$DB_SG" --protocol tcp --port 5432 --source-group "$GW_SG"
On ECS, $ALB_SG goes on the load balancer and $GW_SG on the service. On EKS the AWS Load Balancer Controller makes its own front-end group for the ALB, so those two go unused; an inbound-cidrs annotation restricts the listener instead, and the database group must admit the cluster's security group.
Step 2: IAM roles and the Bedrock use case form
The gateway's task role does one thing: invoke Claude on Bedrock. Per the Bedrock upstream reference, the policy must cover both the cross-region inference profile ARNs and the underlying foundation model ARNs.
cat > bedrock-invoke.json <<EOF
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream", "bedrock:CountTokens"],
"Resource": [
"arn:aws:bedrock:${AWS_REGION}:${ACCOUNT_ID}:inference-profile/us.anthropic.*",
"arn:aws:bedrock:*::foundation-model/anthropic.*"
]
}]
}
EOF
cat > ecs-trust.json <<'EOF'
{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":{"Service":"ecs-tasks.amazonaws.com"},"Action":"sts:AssumeRole"}]}
EOF
aws iam create-role --role-name ccgw-task --assume-role-policy-document file://ecs-trust.json
aws iam put-role-policy --role-name ccgw-task --policy-name bedrock-invoke --policy-document file://bedrock-invoke.json
ECS also needs an execution role, used by the ECS agent (not the gateway) to pull from ECR and inject secrets:
aws iam create-role --role-name ccgw-exec --assume-role-policy-document file://ecs-trust.json
aws iam attach-role-policy --role-name ccgw-exec \
--policy-arn arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy
cat > secrets-read.json <<EOF
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret"],
"Resource": [
"arn:aws:secretsmanager:${AWS_REGION}:${ACCOUNT_ID}:secret:gateway-jwt-secret-??????",
"arn:aws:secretsmanager:${AWS_REGION}:${ACCOUNT_ID}:secret:gateway-oidc-client-secret-??????",
"arn:aws:secretsmanager:${AWS_REGION}:${ACCOUNT_ID}:secret:gateway-postgres-url-??????"
]
}]
}
EOF
aws iam put-role-policy --role-name ccgw-exec --policy-name read-gateway-secrets --policy-document file://secrets-read.json
Why -?????? and not -*? Secrets Manager appends exactly six random characters to every secret ARN, so six ? match that suffix precisely. A -* glob would also match gateway-postgres-url-staging, and a bare gateway-* would grab unrelated secrets in a shared account.
Bedrock enables model access by default in commercial regions. The one remaining gate is Anthropic's one-time use case form: if nobody in the account has submitted it, pick an Anthropic model in the Bedrock console's Model catalog and complete it. Access is immediate. Amazon Bedrock covers the AWS Organizations variant and the permissions the submitter needs.
The EKS track reuses both policy documents on an IRSA role instead of these ECS roles (step 7).
Step 3: RDS for PostgreSQL
Postgres 16, private subnets, no public address, encrypted storage, and TLS forced on the server side:
aws rds create-db-subnet-group --db-subnet-group-name ccgw-db \
--db-subnet-group-description "Claude gateway" --subnet-ids $PRIVATE_SUBNETS
PG_MAJOR=16
aws rds create-db-parameter-group --db-parameter-group-name ccgw-db \
--db-parameter-group-family "postgres${PG_MAJOR}" --description "Force TLS"
aws rds modify-db-parameter-group --db-parameter-group-name ccgw-db \
--parameters "ParameterName=rds.force_ssl,ParameterValue=1,ApplyMethod=immediate"
DB_PASSWORD="$(openssl rand -hex 24)"
aws rds create-db-instance --db-instance-identifier ccgw-db \
--engine postgres --engine-version "$PG_MAJOR" --db-instance-class db.t4g.micro \
--allocated-storage 20 --db-name claude_gateway \
--master-username gateway --master-user-password "$DB_PASSWORD" \
--db-subnet-group-name ccgw-db --db-parameter-group-name ccgw-db \
--vpc-security-group-ids "$DB_SG" --no-publicly-accessible --storage-encrypted
The parameter group family must match the engine's major version, hence the single PG_MAJOR variable.
Warning: A literal
--master-user-passwordshows up in the process table and in audit or EDR logs while the command runs. On shared or monitored hosts, pass it with--cli-input-jsonfrom a file with mode0600.
Once it is up (several minutes), build the connection string:
aws rds wait db-instance-available --db-instance-identifier ccgw-db
DB_HOST="$(aws rds describe-db-instances --db-instance-identifier ccgw-db \
--query 'DBInstances[0].Endpoint.Address' --output text)"
GATEWAY_POSTGRES_URL="postgres://gateway:${DB_PASSWORD}@${DB_HOST}:5432/claude_gateway?sslmode=verify-full"
sslmode=verify-full checks the RDS certificate chain and hostname, not just encryption. The trust anchor is AWS's RDS global certificate bundle, which step 6 bakes into the image and trusts through NODE_EXTRA_CA_CERTS. Do not add sslrootcert= to the URL: the gateway's driver only reads sslmode from the query string and would pass sslrootcert to Postgres as a startup parameter, which the server rejects.
The service or pods must run in this VPC, and only the gateway's group can reach 5432.
Step 4: gateway.yaml
The upstream uses auth: {}, which means the AWS default credential chain: the task role on ECS, the IRSA role on EKS. Two listen fields describe the front end:
public_urlis the externalhttps://origin, required for any non-loopback bind. The gateway builds the IdPredirect_uriand its discovery document from this alone, never fromX-Forwarded-*headers.trusted_proxieslists the ALB's source ranges. An ALB's nodes take addresses from the subnets it is attached to, so use those subnets' CIDRs. That trusts every host in them, so keep your corporate ingress CIDR from overlapping and do not share the subnets with untrusted workloads that could forgeX-Forwarded-For.
The ALB's routing.http.xff_client_port.enabled attribute can be either value; the gateway strips the port from 203.0.113.7:54321 or [2001:db8::1]:54321 forms.
listen:
host: 0.0.0.0
port: 8080
public_url: https://claude.corp.northwind.internal
trusted_proxies: [10.20.1.0/24, 10.20.2.0/24]
oidc:
issuer: https://northwind.okta.com
client_id: 0oa9zzexample7
client_secret: ${OIDC_CLIENT_SECRET} # EKS: ${file:/secrets/oidc-client-secret}
allowed_email_domains: [northwind.co.uk]
userinfo_fallback: true # Okta org server omits email and groups
scopes: [openid, profile, email, offline_access, groups]
session:
jwt_secret: ${GATEWAY_JWT_SECRET} # EKS: ${file:/secrets/jwt-secret}
ttl_hours: 8 # lower towards 1 for faster revocation
store:
postgres_url: ${GATEWAY_POSTGRES_URL} # EKS: ${file:/secrets/postgres-url}
# readiness_grace_seconds: 300 # ride out an RDS failover
upstreams:
- provider: bedrock
region: us-west-2 # keep in step with $AWS_REGION
auth: {}
Only oidc is Okta-specific. For Microsoft Entra ID, use issuer https://login.microsoftonline.com/<tenant-id>/v2.0, drop userinfo_fallback and the groups scope, and remember Entra sends group GUIDs, so managed.policies must match GUIDs, or App Roles with oidc.groups_claim: roles.
Step 5: secrets
aws secretsmanager create-secret --name gateway-jwt-secret --secret-string "$(openssl rand -base64 32)"
aws secretsmanager create-secret --name gateway-oidc-client-secret --secret-string file://okta-secret.txt
aws secretsmanager create-secret --name gateway-postgres-url --secret-string "$GATEWAY_POSTGRES_URL"
Keep the ARNs each call prints; the task definition references them. The same process-table caveat applies to literal --secret-string values, so prefer file:// with a 0600 file, as in the middle line.
gateway.yaml itself contains no secrets because every credential is expanded at boot (see keeping secrets out of the file). Delivery differs:
- ECS: the image build copies
gateway.yamlto/etc/claude/gateway.yaml, and the task definition injects the three secrets as environment variables. - EKS: mount
gateway.yamlfrom a ConfigMap and the secrets as files under/secrets. Source the Kubernetes Secrets from Secrets Manager with External Secrets Operator or the Secrets Store CSI driver's AWS provider, or create them withkubectl.
Step 6: build and push to ECR
Build per the image requirements, with the linux-x64 glibc binary at ./claude. On ECS, the Dockerfile copies the filled-in gateway.yaml into the image, which is why the build follows step 4; on EKS that copy goes unused.
Add the RDS CA bundle. AWS appends new regional CAs over time, so download it on every build rather than committing it or pinning a checksum:
curl -fL --proto '=https' -o rds-global-bundle.pem \
https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem
Two lines in the Dockerfile trust it:
COPY rds-global-bundle.pem /etc/claude/rds-global-bundle.pem
ENV NODE_EXTRA_CA_CERTS=/etc/claude/rds-global-bundle.pem
Create an ECR repository with immutable tags, so a pinned tag can never be quietly re-pointed, and push:
aws ecr create-repository --repository-name claude-gateway \
--image-tag-mutability IMMUTABLE --image-scanning-configuration scanOnPush=true
aws ecr get-login-password --region "$AWS_REGION" \
| docker login --username AWS --password-stdin "${ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com"
IMAGE="${ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/claude-gateway:2.1.290-1"
docker build --platform=linux/amd64 -t "$IMAGE" . && docker push "$IMAGE"
For Graviton, build linux/arm64 with the linux-arm64 binary and set cpuArchitecture to ARM64 in the task definition.
Step 7: deploy
Option A: ECS Fargate
Cluster and a log group for stderr (audit events and operational logs). Without a retention policy CloudWatch keeps logs forever; match it to your audit policy:
aws ecs create-cluster --cluster-name claude-gateway
aws logs create-log-group --log-group-name /ecs/claude-gateway
aws logs put-retention-policy --log-group-name /ecs/claude-gateway --retention-in-days 180
Task definition, with the task role for Bedrock and the execution role for secrets:
{
"family": "claude-gateway",
"networkMode": "awsvpc",
"requiresCompatibilities": ["FARGATE"],
"cpu": "1024",
"memory": "2048",
"runtimePlatform": { "cpuArchitecture": "X86_64", "operatingSystemFamily": "LINUX" },
"executionRoleArn": "arn:aws:iam::111122223333:role/ccgw-exec",
"taskRoleArn": "arn:aws:iam::111122223333:role/ccgw-task",
"containerDefinitions": [{
"name": "gateway",
"image": "111122223333.dkr.ecr.us-west-2.amazonaws.com/claude-gateway:2.1.290-1",
"portMappings": [{ "containerPort": 8080 }],
"secrets": [
{ "name": "GATEWAY_JWT_SECRET", "valueFrom": "arn:aws:secretsmanager:...:gateway-jwt-secret-AbC123" },
{ "name": "OIDC_CLIENT_SECRET", "valueFrom": "arn:aws:secretsmanager:...:gateway-oidc-client-secret-DeF456" },
{ "name": "GATEWAY_POSTGRES_URL", "valueFrom": "arn:aws:secretsmanager:...:gateway-postgres-url-GhI789" }
],
"logConfiguration": {
"logDriver": "awslogs",
"options": { "awslogs-group": "/ecs/claude-gateway", "awslogs-region": "us-west-2", "awslogs-stream-prefix": "gw" }
}
}]
}
aws ecs register-task-definition --cli-input-json file://ccgw-task.json
Internal ALB and target group. --ip-address-type ipv4 is essential: an internal dual-stack ALB publishes public-range AAAA records, which the /login private-network check rejects.
ALB_ARN="$(aws elbv2 create-load-balancer --name claude-gateway --scheme internal \
--type application --ip-address-type ipv4 --subnets $PRIVATE_SUBNETS \
--security-groups "$ALB_SG" --query 'LoadBalancers[0].LoadBalancerArn' --output text)"
TG_ARN="$(aws elbv2 create-target-group --name claude-gateway --protocol HTTP --port 8080 \
--vpc-id "$VPC_ID" --target-type ip --health-check-path /readyz \
--query 'TargetGroups[0].TargetGroupArn' --output text)"
HTTPS listener with a modern TLS policy (omitting --ssl-policy falls back to ELBSecurityPolicy-2016-08, which still allows TLS 1.0 and 1.1), plus a longer idle timeout. The ALB's default is 60 seconds of silence; the gateway's keepalive pings fit inside that, but an hour adds margin:
aws elbv2 create-listener --load-balancer-arn "$ALB_ARN" --protocol HTTPS --port 443 \
--ssl-policy ELBSecurityPolicy-TLS13-1-2-2021-06 \
--certificates CertificateArn=arn:aws:acm:us-west-2:111122223333:certificate/example \
--default-actions Type=forward,TargetGroupArn="$TG_ARN"
aws elbv2 modify-load-balancer-attributes --load-balancer-arn "$ALB_ARN" \
--attributes Key=idle_timeout.timeout_seconds,Value=3600
Service, with the deployment circuit breaker so a bad image or unbootable config rolls back instead of looping:
aws ecs create-service --cluster claude-gateway --service-name claude-gateway \
--task-definition claude-gateway --desired-count 2 --launch-type FARGATE \
--deployment-configuration "deploymentCircuitBreaker={enable=true,rollback=true}" \
--health-check-grace-period-seconds 60 \
--network-configuration "awsvpcConfiguration={subnets=[$(echo $PRIVATE_SUBNETS | tr ' ' ',')],securityGroups=[$GW_SG],assignPublicIp=DISABLED}" \
--load-balancers "targetGroupArn=$TG_ARN,containerName=gateway,containerPort=8080"
The 60-second grace lets a cold task pull, connect to the store and answer its first health check. The /readyz check means a task that cannot reach Postgres never enters rotation; to keep tasks in through an RDS failover, set store.readiness_grace_seconds as described under when dependencies fail.
All egress (Bedrock, IdP, Secrets Manager, ECR, CloudWatch Logs) goes through the NAT gateway because tasks have no public IP. To keep Bedrock traffic off the public path, create a bedrock-runtime interface VPC endpoint and point the upstream's base_url at it; the IdP still needs internet egress.
Finally, alias your own internal hostname to the ALB in the private hosted zone and use it as listen.public_url. The ALB's *.elb.amazonaws.com name resolves privately on an internal ALB but cannot carry your ACM certificate. Update the Okta redirect URI to <public_url>/oauth/callback before the first sign-in. Because the config is baked into the image on ECS, any change to public_url means rebuild, push under a new tag, register a new task definition revision and redeploy.
Option B: EKS
You need kubectl, eksctl, and an existing cluster in $VPC_ID with an IAM OIDC provider and the AWS Load Balancer Controller. The database group must admit the cluster's pod or node security group instead of $GW_SG.
On EKS, credentials come from IRSA, which needs a role whose trust policy federates on the cluster's OIDC provider for system:serviceaccount:claude-gateway:gateway. eksctl create iamserviceaccount creates it, attaches policies and annotates the service account in one go:
BEDROCK_POLICY_ARN="$(aws iam create-policy --policy-name ccgw-bedrock-invoke \
--policy-document file://bedrock-invoke.json --query Policy.Arn --output text)"
SECRETS_POLICY_ARN="$(aws iam create-policy --policy-name ccgw-secrets-read \
--policy-document file://secrets-read.json --query Policy.Arn --output text)"
kubectl create namespace claude-gateway
eksctl create iamserviceaccount --cluster platform-prod --region "$AWS_REGION" \
--namespace claude-gateway --name gateway --role-name ccgw-irsa \
--attach-policy-arn "$BEDROCK_POLICY_ARN" --attach-policy-arn "$SECRETS_POLICY_ARN" --approve
The secrets policy is only needed when pods read Secrets Manager themselves, as the CSI driver's AWS provider does. Keep both actions: the provider calls DescribeSecret when reconciling rotations, so a GetSecretValue-only grant mounts once and then misses rotated values.
Deploy a Deployment, Service and Ingress as in the Kubernetes section, with serviceAccountName: gateway, the ConfigMap and /secrets mounts, and readiness on /readyz. Annotate the Ingress:
metadata:
annotations:
alb.ingress.kubernetes.io/scheme: internal
alb.ingress.kubernetes.io/target-type: ip
alb.ingress.kubernetes.io/ip-address-type: ipv4
alb.ingress.kubernetes.io/inbound-cidrs: 10.40.0.0/14
alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:us-west-2:111122223333:certificate/example
alb.ingress.kubernetes.io/ssl-policy: ELBSecurityPolicy-TLS13-1-2-2021-06
alb.ingress.kubernetes.io/load-balancer-attributes: idle_timeout.timeout_seconds=3600
inbound-cidrs replaces the controller's 0.0.0.0/0 default on its front-end group. With IRSA the SDK exchanges a projected service account token with STS, so pods never need instance metadata and an egress NetworkPolicy can block 169.254.169.254.
Step 8: point laptops at it
Developers cannot reach the gateway from /login until forceLoginMethod and forceLoginGatewayUrl are on their machines via MDM-delivered managed settings; there is no manual option in the picker. See point machines at the gateway.
The companion bundle
The Claude Code GitHub repository has an examples/gateway/aws bundle that packages this walkthrough:
setup.shruns the sameawscommands for the ECS track. It is idempotent and every default can be overridden by environment variable. You still create the Okta client secret and the ACM certificate; without them it skips the ECS and ALB steps, names what is missing and prints thecreate-secretcommand. The use case form and Route 53 alias are printed as next steps, and the MDM push is manual.gateway.yaml.exampleis the step 4 template with optional keys commented out; replace everyREPLACE_ME.Dockerfilebuilds from thelinux-x64binary and embeds yourgateway.yamland the RDS bundle.setup.shonly downloads the bundle if missing, so delete it and rebuild under a new tag after an AWS CA rotation. It tags images with a hash of the config file, so any config edit is a new tag.terraform/provisions the same ECS scope declaratively, with the VPC and subnets as inputs. Terraform creates the repository but does not build the image, so apply in two passes: a targeted apply for ECR, build and push, then the full apply. Its README covers variables, remote state and teardown.
Like this page, it is an example to adapt, not a supported product.
AWS-specific troubleshooting
The general table is in the deployment guide.
| Symptom | Cause | Fix |
|---|---|---|
/login: Gateway hosts must be on your organization's private network; ... resolves to the public (or unrecognized) address ... | Dual-stack internal ALB publishing public-range AAAA records | Create the ALB with --ip-address-type ipv4, or use an internal-only name without AAAA |
Every Bedrock request 502s, log says Could not load credentials from any providers | ECS on EC2 without a task role, or EKS without IRSA, so credentials come from instance metadata, which IMDSv2's default hop limit of 1 blocks in containers. Fargate task roles and IRSA are unaffected | Use task roles or IRSA; otherwise raise the hop limit to 2 with aws ec2 modify-instance-metadata-options |
Bedrock 403 AccessDeniedException | Use case form not submitted, first-invoke Marketplace subscription still completing, or the policy lacks one of the two ARN families | Submit the form; retry after a few minutes on a first invoke; grant both ARN families |
ValidationException about on-demand throughput | A custom models entry uses a bare foundation model ID the region only serves via inference profiles | Map to the us.anthropic.* profile ID; the built-in catalogue already does |
ECS task stops with ResourceInitializationError before any gateway log | Execution role cannot read secrets, or no route to Secrets Manager or ECR | Grant secretsmanager:GetSecretValue on the three ARNs; provide NAT egress or interface endpoints for Secrets Manager, ECR and CloudWatch Logs plus an S3 gateway endpoint |
| Postgres connection timeout at boot | Database group does not admit the gateway's group on 5432, or wrong VPC | Fix the rule; run in the database's VPC |
| Postgres TLS verification error at boot | Image does not trust the RDS bundle | Add the two Dockerfile lines, rebuild under a new tag, redeploy |
| Streams drop after a quiet spell | Gateway older than v2.1.229 sends nothing during quiet periods on Bedrock or Claude Platform on AWS upstreams, and the ALB cuts at 60 seconds. Newer gateways ping after about 15 seconds of silence | Upgrade to v2.1.229+, or set idle_timeout.timeout_seconds to 3600 |
Telemetry on AWS
Claude Code emits OTLP metrics, logs and opt-in traces (Monitoring usage). In sessions signed in through /login, the CLI stamps every export with user.id, user.email and user.groups from the IdP, so usage rolls up per developer with no per-laptop OTEL setup.
The gateway acts as an authenticated OTLP relay. Set telemetry.forward_to together with listen.public_url, and it pushes exporter settings to connected clients and forwards their OTLP traffic unchanged to each listed destination. Each destination opts into metrics, logs and traces separately (metrics only by default). Nothing is buffered or stored at the gateway. Client telemetry is off until you configure this, and each interactive client shows a security approval dialog for the pushed settings.
On AWS:
- Client signals: forward to an OpenTelemetry collector such as the AWS Distro for OpenTelemetry collector, run as its own internal
https://service, and export to CloudWatch, Amazon Managed Service for Prometheus or any OTLP backend. - Gateway logs: on Fargate the
awslogsdriver already ships stderr to/ecs/claude-gateway. On EKS, pod logs do not reach CloudWatch by default, so install the CloudWatch Observability add-on with container logs, or a Fluent Bit DaemonSet, or you lose the audit trail. Query with CloudWatch Logs Insights and alarm on metric filters. - Container metrics:
aws ecs update-cluster-settings --cluster claude-gateway --settings name=containerInsights,value=enabledon ECS; the CloudWatch Observability add-on on EKS. - Spend: telemetry is after the fact; spend limits are the live view and the enforcement.
Splitting the Bedrock bill
By default AWS sees all gateway spend under one principal, the task or IRSA role. Two techniques split it in AWS's own billing data, and they can be combined.
Per developer: assume a role per person
Create a second role that holds the Bedrock permissions and trusts the gateway's principal, allow the gateway's principal sts:AssumeRole on it, and set assume_role with session_name: email on the Bedrock upstream. The gateway then assumes the role once per developer per hour with their email as the session name and signs their requests with those credentials. It needs a gateway on v2.1.281 or later, and the role can live in another account (Bedrock in another AWS account).
In Terraform, alongside the task role:
resource "aws_iam_role" "per_dev_bedrock" {
name = "ccgw-per-developer"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{ Effect = "Allow", Action = "sts:AssumeRole", Principal = { AWS = aws_iam_role.gateway_task.arn } }]
})
}
resource "aws_iam_role_policy" "per_dev_invoke" {
role = aws_iam_role.per_dev_bedrock.id
policy = file("${path.module}/bedrock-invoke.json")
}
resource "aws_iam_role_policy" "gateway_can_assume" {
role = aws_iam_role.gateway_task.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [{ Effect = "Allow", Action = "sts:AssumeRole", Resource = aws_iam_role.per_dev_bedrock.arn }]
})
}
With assume_role on, every Bedrock call from that upstream, including the free CountTokens used for spend metering, is signed with the assumed credentials, so the gateway's own principal only needs Bedrock rights for upstreams without it. The subnets need a route to sts.<region>.amazonaws.com (the NAT gateway, or an STS interface endpoint), and each active developer costs one STS call per hour per replica.
Requests then appear as arn:aws:sts::<account>:assumed-role/<role>/<email>. To report spend per principal, enable AWS's IAM principal cost allocation and use a billing export that includes principal data.
Per team: application inference profiles
This needs only the models and managed sections. Create a Bedrock application inference profile per team and model, tag each with the team, and activate that tag for cost allocation. Then give each team its own model ID and pin each IdP group to it:
models:
- id: payments-claude-sonnet-4-6
upstream_model:
bedrock: arn:aws:bedrock:us-west-2:111122223333:application-inference-profile/pay01
- id: search-claude-sonnet-4-6
upstream_model:
bedrock: arn:aws:bedrock:us-west-2:111122223333:application-inference-profile/srch01
managed:
policies:
- match: {groups: [eng-payments]}
cli: {availableModels: [payments-claude-sonnet-4-6], enforceAvailableModels: true}
- match: {groups: [eng-search]}
cli: {availableModels: [search-claude-sonnet-4-6], enforceAvailableModels: true}
- match: {}
cli: {availableModels: [claude-sonnet-4-6, claude-haiku-4-5], enforceAvailableModels: true}
Developers in a pinned team start Claude Code with --model payments-claude-sonnet-4-6 (their own team's ID), because a session on the default model is refused for them. The gateway enforces availableModels on every request, not just in the picker, and AWS billing groups spend by the activated tag. Keep the match: {} catch-all: without it, someone matching no policy gets the whole catalogue and could bill either team's profile.
The trade-offs: config grows as teams times models, and whichever role signs the upstream (the gateway's principal, or the assumed role) must be allowed to invoke application-inference-profile/* ARNs. pricing explains how the gateway's own spend meter prices custom IDs. AWS also maintains its own Claude apps gateway samples in the aws-samples/anthropic-on-aws GitHub repository.