Skip to content

Claude apps gateway spend limits

Give each developer, group or the whole organisation a daily, weekly or monthly spend ceiling that the Claude apps gateway enforces on every request.

A Claude apps gateway sends everybody's inference through one shared upstream credential. Your cloud bill therefore shows one big number against one principal, and nothing stops a single runaway agent loop from eating a month's commitment overnight. Spend limits fix that. You set a cap per developer, per IdP group or as an organisation default, and once someone passes it the gateway answers their next request with a 429 until the period resets or an admin raises the cap.

This page covers turning the feature on, setting caps, how the gateway prices requests, what developers see in Claude Code, the admin API and how long the data is kept. It assumes you already have a gateway running from Claude apps gateway.

Tip: On Amazon Bedrock you can also get per-developer attribution in the AWS bill itself with assume_role and session_name. See per-developer AWS cost attribution. Spend limits are the gateway's own view and circuit breaker on top of whatever the provider bills.

Turning it on

Spend enforcement runs only when gateway.yaml contains an admin block. That block defines who may call the admin API, and its presence switches on live checks on every /v1/messages request.

admin:
  write_keys:
    - { id: finops-terraform, key: "${SPEND_ADMIN_KEY_TF}" }
  read_keys:
    - { id: grafana-export, key: "${SPEND_READ_KEY}" }
  admin_groups: [eng-leads]
  blocked_message: "Ask in #claude-budget on Slack for a top-up."

The caps themselves are not in YAML. You manage them through the API, which means changes take effect immediately without a redeploy.

Authenticating

There are two ways to call the admin endpoints:

MethodAccessShows in audit as
x-api-key header matching an entry in admin.write_keysRead and writeadmin-key:<id>
x-api-key header matching an entry in admin.read_keysGET onlyadmin-key:<id>
A gateway bearer token whose groups claim includes one of admin.admin_groupsRead and writeoidc:<sub>

Give each automation its own key ID so the audit trail tells you which one did what. For people, prefer the group route so changes are attributed to a real identity.

Setting caps

Each POST /v1/organizations/spend_limits creates or replaces one cap identified by its scope and period. Amounts are whole US cents as a string.

Caps are always in US dollars, even if your finance team thinks in pounds. A sensible first move is a $300 monthly default for everyone:

curl -sS https://claude.tools.northwind.internal/v1/organizations/spend_limits \
  -H "x-api-key: $SPEND_ADMIN_KEY_TF" \
  -H "Content-Type: application/json" \
  -d '{"scope":{"type":"organization"},"amount":"30000","period":"monthly"}'

Then a weekly ceiling of $40 for the interns group:

curl -sS https://claude.tools.northwind.internal/v1/organizations/spend_limits \
  -H "x-api-key: $SPEND_ADMIN_KEY_TF" \
  -H "Content-Type: application/json" \
  -d '{"scope":{"type":"rbac_group","rbac_group_id":"summer-interns"},"amount":"4000","period":"weekly"}'

And a personal exception for one developer running a large migration, keyed on their OIDC sub:

curl -sS https://claude.tools.northwind.internal/v1/organizations/spend_limits \
  -H "x-api-key: $SPEND_ADMIN_KEY_TF" \
  -H "Content-Type: application/json" \
  -d '{"scope":{"type":"user","user_id":"00u8k2migration"},"amount":"150000","period":"monthly"}'

Request fields

FieldAccepted valuesMeaning
scope.typeuser, rbac_group, organizationWho the cap applies to.
scope.user_idOIDC subRequired for user. The stable ID your IdP assigns, not the email.
scope.rbac_group_idIdP group nameRequired for rbac_group. Same names you match in managed policies.
amountString of whole USD cents, or nullnull means unlimited. "0" blocks every request.
perioddaily, weekly, monthlyOne cap per period per scope. They are independent: exceeding any one blocks the developer.

The API mirrors the wire shapes of Anthropic's public Admin API spend-limit endpoints, so a client built for that contract can point at the gateway by changing its base URL. One difference: the gateway accepts all three scope types, whereas the public POST currently accepts user scope only.

How a developer's cap is resolved

Group and organisation caps are per-seat defaults, not shared pools. Every member of summer-interns gets their own $40 a week.

For each period, the gateway works out a developer's effective cap in this order:

  1. A cap set directly on that user.
  2. Otherwise, the most restrictive cap among the groups they belong to.
  3. Otherwise, the organisation default.
  4. Otherwise, unlimited.

If you would rather a developer in several groups get the most generous group cap, set admin.group_limit_mode: max. The default is min.

What enforcement looks like

On each /v1/messages request the gateway fetches the developer's caps and period-to-date spend in a single Postgres query. If they are over any cap, the response is:

  • HTTP 429
  • error.type: billing_error
  • header x-should-retry: false
  • header retry-after with the seconds until the cap resets
  • a message naming the period and reset time, such as spend limit reached (weekly; resets 2026-10-12 00:00 UTC), followed by your admin.blocked_message

If several caps are exceeded, the message names whichever resets last. Gateways older than v2.1.225 sent only spend limit reached with no period, reset time or retry-after. On v2.1.227 or later, <public_url>/protocol documents the exact headers and body.

Periods reset on UTC calendar boundaries: daily at 00:00 UTC, weekly on Monday, monthly on the 1st. /v1/messages/count_tokens is never blocked because counting tokens costs nothing.

Pricing each request

After a response finishes, a meter reads the token usage and adds the cost to the daily, weekly and monthly counters. It never touches the bytes going to the client, so a metering fault cannot break a response. Treat the figures as a USD estimate for circuit-breaking, not an invoice. Reconcile real billing against your provider.

The meter chooses a rate in this order:

  1. A matching row in pricing.overrides for the upstream that served the request (gateway v2.1.227 or later).
  2. List price for the upstream model ID, if Claude Code's cost table recognises it. Anthropic, Bedrock, Agent Platform and Foundry ID forms are all understood.
  3. List price for the models[].id you mapped to that upstream ID. This covers strings with no model name in them, such as a Bedrock application inference profile ARN or a Foundry deployment name (v2.1.218 or later).
  4. A fallback "unknown model" tier of $5 input and $25 output per million tokens, so nothing is ever free. The gateway warns at boot and once per unknown ID at runtime.

Whatever rate is chosen is then multiplied by pricing.multiplier, which defaults to 1.

Aborted requests still count. If a stream ends without the upstream's final usage frame, the meter bills a floor of roughly four characters per output token for text already delivered, so cancelling early is not a way round a cap.

When Postgres is down

The pre-request check has a two-second timeout. By default enforcement fails open: if the database is slow or unreachable, the request goes ahead, a warning is logged and the response omits the anthropic-ratelimit-unified-* headers.

Set enforcement.fail_closed_on_error: true to fail closed instead. Requests then get the same 429 billing_error but with the message spend limit unavailable and no reset time or retry-after.

Choose based on what hurts more. Fail-open means a database blip does not become an inference outage. Fail-closed guarantees no unmetered spend at the price of blocking everyone during an outage. Either way, fail-open only helps while your load balancer still routes to the gateway; store.readiness_grace_seconds keeps replicas reporting ready through a short database failover.

What developers see in Claude Code

Developers are not left guessing. Claude Code warns once their most-used cap passes 75% and again past 95%. When a request is blocked it shows the gateway's message verbatim, including your blocked_message. The cap also appears in /usage and is passed to the status line script.

DisplayMinimum on developer machineMinimum on gateway
75% and 95% warningsv2.1.225v2.1.225
Spend limit bar in /usage with percentage and reset time, plus rate_limits.spend_limit in status line inputv2.1.251v2.1.225
Dollar amounts in the bar (for example "$212.40 / $300.00 spent this month") and amounts plus period in status line inputv2.1.284v2.1.284

The warnings come from anthropic-ratelimit-unified-* headers the gateway adds to successful responses for anyone with a cap. Those headers always describe the developer's own cap; the gateway strips the provider's rate-limit headers, which describe your shared quota.

The dollar figures need a separate request from Claude Code to the gateway. If you set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC on developer machines, that request is skipped and the display stays percentage-only. The same happens with a gateway older than v2.1.284.

I find the status line display is what actually changes behaviour. A developer who can see "71% of weekly" in the corner of the terminal tends to switch to a smaller model for routine work without being asked.

Admin API reference

All paths sit under /v1/organizations/spend_limits.

Method and pathWhat it does
GET /v1/organizations/spend_limitsList caps. Query: limit, after_id, before_id, scope_type (organization, rbac_group or user).
POST /v1/organizations/spend_limitsCreate or replace the cap for a {scope, period} pair.
GET /v1/organizations/spend_limits/{id}Fetch one cap by its spl_ ID.
DELETE /v1/organizations/spend_limits/{id}Remove a cap. Returns {type: "spend_limit_deleted", id}.
GET /v1/organizations/spend_limits/effectiveResolved cap and period-to-date spend for each principal and period.
GET /v1/organizations/spend_limits/auditChange history, newest first. Query: limit, after_id.

Conventions follow Anthropic's Admin API: every object has a type, IDs are prefixed spl_, amounts are cent strings (a POST with any non-USD currency gets a 400), errors use the {type: "error", error: {type, message}, request_id} envelope, and every admin response carries a request-id header.

Every change writes a before-and-after row to the admin_audit table in the same transaction. Only the spend-limit endpoints exist; other Admin API surfaces, such as spend_limit_increase_requests, are not served.

The effective view

/effective returns Anthropic's SpendSummary shape, one row per principal per period, with the resolved cap, spend so far and an actor object. Gateway-specific details:

  • user_id is the OIDC sub.
  • actor.name and actor.email_address stay null until that person's first request through the gateway, because the gateway has no directory and learns these from each session token.
  • Each row adds a groups array of last-seen IdP groups so a dashboard can show every tier that applies. Clients built for Anthropic's shape simply ignore it.
  • Without a user_ids[] filter, only principals with recorded spend appear. The gateway cannot list members who have never used it.

Group caps are resolved against those last-seen groups using the same group_limit_mode, so what you see matches what is enforced.

ParameterUse
user_ids[]Repeatable. Restrict to specific OIDC sub values.
period[]Repeatable. daily, weekly or monthly.
sortspend_desc for top spenders first. Needs exactly one period[].
qCase-insensitive substring match across sub, last-seen email and display name.
limit / pagePage size 1 to 1000 (default 20) and the opaque next_page cursor.

A typical "who is burning the budget this week" query:

curl -sS -G https://claude.tools.northwind.internal/v1/organizations/spend_limits/effective \
  -H "x-api-key: $SPEND_READ_KEY" \
  --data-urlencode "period[]=weekly" \
  --data-urlencode "sort=spend_desc" \
  --data-urlencode "limit=10"

Warning: q and user_ids[] travel in the query string, so proxies and load balancers in front of the gateway will capture them in access logs. Scrub those parameters there if your PII policy is strict.

Paging

  • The plain list pages with after_id or before_id (mutually exclusive spl_ IDs). Results are ordered by creation and has_more reflects the direction.
  • /effective pages with the opaque next_page value passed back as ?page=. Principals are sorted ascending so pages stay stable while spend accrues.
  • /audit pages with after_id, the numeric id of the last event you saw. Its default limit is 100 and has_more is exact.
  • limit is 1 to 1000 everywhere, defaulting to 20 except on /audit.

Data kept and for how long

An hourly sweep applies retention to four tables:

TableHoldsKept for
spendPeriod-to-date counters per principal, in centsadmin.spend_retention_months, default 13
spend_limitsThe caps you configuredUntil you delete them
admin_auditChange historyadmin.audit_retention_days, default 365
principal_emailsLast-seen email, display name and groups (personal data)admin.identity_retention_days after last activity, default 90

When someone leaves, delete any personal cap with DELETE /v1/organizations/spend_limits/{id} and let the rest age out. For an immediate erasure, for example a data subject access request, delete their row directly:

DELETE FROM principal_emails WHERE principal = '00u8k2migration';

That is the only table holding name, email and groups. The spend and admin_audit rows refer only to the pseudonymous sub and expire on their own schedule. Back up the database, because with spend limits on it holds durable financial and audit data; see the deployment guide.