Skip to content

Google Vertex AI

Run Claude Code on Claude models in your Google Cloud project via Agent Platform (formerly Vertex AI): wizard, credentials, regions, model pinning, IAM and fixes.

Google now calls the service Agent Platform, but most people (and the Claude Code login menu) still say Vertex AI, so I use both names here. Running Claude Code through it keeps inference inside your Google Cloud project, billed through GCP and controlled by IAM and Cloud Audit Logs. As with the other cloud providers, claude.ai-only features are unavailable; see feature availability.

Individuals can use the login wizard. For team rollouts and CI, use the manual setup and pin model versions.

Before you start

  • A GCP account with billing enabled.
  • A project with the Agent Platform API enabled.
  • Access to the Claude models you want, granted in Model Garden.
  • The gcloud CLI installed and configured.
  • Quota in the region or location you plan to use.

The quick route: the login wizard

  1. Prepare the project once. Enable the API (see step 1 below) and request access to Claude models in Model Garden.
  2. Open the wizard. Run claude, choose 3rd-party platform, then Google Vertex AI. If you are already signed in, /login gets you to the same menu.
  3. Answer the prompts. Choose Application Default Credentials from gcloud, a service account key file, or credentials already in the environment; then the project and region. The wizard checks which models the project can call and lets you pin them.

Results go into the env block of ~/.claude/settings.json (or $CLAUDE_CONFIG_DIR/settings.json). Rerun with /setup-vertex to change credentials, project, region or pins.

Regions and locations

CLOUD_ML_REGION accepts three forms:

FormExampleNotes
Global endpointglobalBest availability, but not every model supports it
Multi-region locationeu, usUses aiplatform.eu.rep.googleapis.com / aiplatform.us.rep.googleapis.com
Specific regioneurope-west1, us-east5Model availability varies region to region

Claude Code picks the right hostname for each. Not every default model exists on every endpoint type, so you may need to change location or choose a different model.

The manual route

1. Enable the API

gcloud config set project acme-claude-dev
gcloud services enable aiplatform.googleapis.com

2. Request model access

In Model Garden, search for Claude, request access to the models you need and wait for approval, which can take a day or two.

3. Provide credentials

Claude Code uses standard Google Cloud authentication through Application Default Credentials. That includes X.509 certificate-based Workload Identity Federation: point GOOGLE_APPLICATION_CREDENTIALS at your credential configuration file.

Note: Requests always go to the project in ANTHROPIC_VERTEX_PROJECT_ID, even if GCLOUD_PROJECT, GOOGLE_CLOUD_PROJECT or your credentials file names a different one. This catches people out when their default gcloud project is something else.

Refreshing expired credentials

gcpAuthRefresh runs a command when credentials have expired or cannot be loaded, then retries the request:

{
  "gcpAuthRefresh": "gcloud auth application-default login",
  "env": { "ANTHROPIC_VERTEX_PROJECT_ID": "acme-claude-dev" }
}

Before running it, Claude Code tries to get an access token with your current credentials and skips the command if they still work. If that check takes longer than five seconds, it also skips the command and only runs it after a request actually fails (since v2.1.261; before that, a slow check could pop a browser at startup unnecessarily). Output is shown, but the command cannot receive input, so browser-based flows are ideal. It times out after three minutes. A gcpAuthRefresh in project settings is subject to the same workspace trust rule as hooks in settings files, which includes -p runs in folders you have never trusted.

4. Point Claude Code at Vertex AI

export CLAUDE_CODE_USE_VERTEX=1
export CLOUD_ML_REGION=global
export ANTHROPIC_VERTEX_PROJECT_ID=acme-claude-dev

# optional: a custom endpoint or gateway
# export ANTHROPIC_VERTEX_BASE_URL=https://vertex-gw.example.internal

# with a global endpoint, send models that lack global support to a region
export VERTEX_REGION_CLAUDE_HAIKU_4_5=europe-west1
export VERTEX_REGION_CLAUDE_4_6_SONNET=europe-west4

Most model versions have a matching VERTEX_REGION_CLAUDE_* variable; the environment variables reference lists them. Model Garden shows which models support the global endpoint.

Region values that do not look like a region (containing a slash, dot or space) are ignored. A bad VERTEX_REGION_CLAUDE_* falls back to CLOUD_ML_REGION; a bad or missing CLOUD_ML_REGION falls back to us-east5.

Other things to know:

  • Prompt caching is on automatically. DISABLE_PROMPT_CACHING=1 turns it off; ENABLE_PROMPT_CACHING_1H=1 requests a one-hour TTL at a higher write price.
  • Rate limit increases come from Google Cloud support.
  • /logout does nothing here; Google credentials handle auth.
  • MCP tool search is on by default for Opus 4.5, Sonnet 4.5, Haiku 4.5 and later. Earlier models, including all Claude 3.x, load MCP tool definitions upfront because their serving stack rejects the required beta header, and ENABLE_TOOL_SEARCH=true cannot override that. ENABLE_TOOL_SEARCH=false turns tool search off everywhere. (Before v2.1.221 it was off by default on Vertex AI for every model.)

5. Pin your models

Warning: Pin versions before rolling out. Unpinned aliases follow Claude Code's built-in Vertex AI defaults, which may lag the latest release or name models your project has not enabled. Claude Code falls back at startup if a default is missing, but pinning puts you in control of when people move.

Unpinned, opus resolves to Opus 5.5 and sonnet to Sonnet 4.5. To pin:

export ANTHROPIC_DEFAULT_OPUS_MODEL='claude-opus-4-8'
export ANTHROPIC_DEFAULT_SONNET_MODEL='claude-sonnet-5'
export ANTHROPIC_DEFAULT_HAIKU_MODEL='claude-haiku-4-5@20251001'
Default with nothing pinnedModel ID
Primary modelclaude-opus-5-5
Small/fast modelclaude-sonnet-4-5@20250929

Background work such as session titles normally uses a Haiku-class model, but on Vertex AI it uses Sonnet because Haiku is not enabled everywhere. Choosing a primary model (--model, ANTHROPIC_MODEL, the model setting or ANTHROPIC_DEFAULT_MODEL) makes background work use that model too, and so does pinning Opus without pinning Sonnet. Set ANTHROPIC_DEFAULT_HAIKU_MODEL to an enabled model if you want Haiku for it.

Warning: Unpinned deployments have been billed at Opus rates since v2.1.207. To stay on Sonnet 4.5, set ANTHROPIC_MODEL to its full ID. Steering with ANTHROPIC_DEFAULT_SONNET_MODEL alone keeps your Sonnet as the default.

Version history for older fleets: before v2.1.280 the default and opus were Opus 5 (from v2.1.219); v2.1.207 to v2.1.218 used Opus 4.8; before v2.1.207 the default was Sonnet 4.5, opus meant Opus 4.6, and background tasks used the primary model.

6. Check it

Run /status. The API provider line should read Google Vertex AI, and GCP project, Default region and Model should show your values. No provider line means the variables are not reaching the process: export them in the launching shell or put them in settings env.

Startup model checks

At launch, Claude Code confirms the models it intends to use are callable in your project:

  • Old pin, newer version available: you are offered an update. Accepting writes the new ID to user settings and restarts; declining is remembered until the next default change.
  • No pin, default unavailable: a fallback for this session only (earlier versions of the same model, then from Opus to the default Sonnet), with a notice. Enable the newer model in Model Garden or pin to make it stick.
  • Specific version chosen at launch via --model, ANTHROPIC_MODEL or model: treated as that tier's pin, with no check of the default it replaces. Aliases and unrecognised IDs are not pins.

Refusals are remembered for up to a day per machine; a refused current default is retried after ten minutes. CLAUDE_CODE_SKIP_MODEL_ACCESS_MEMORY=1 disables the memory.

With an enforced allowlist (enforceAvailableModels, v2.1.287+) the checks only consider models in availableModels, written as the IDs Claude Code sends:

{
  "availableModels": ["claude-opus-4-8", "claude-sonnet-5"],
  "enforceAvailableModels": true
}

If a model is disabled mid-session, unpinned tiers switch automatically (Switched to <fallback> because <model> is not available) using the same order. Pinned versions fail instead. In auto mode, only models auto mode supports on Vertex AI are candidates. A configured fallback model chain takes priority, and CLAUDE_CODE_DISABLE_MODEL_ACCESS_FALLBACK=1 makes refusals fail rather than switch (remove any fallback chain too if every refusal should fail).

IAM

roles/aiplatform.user covers what Claude Code needs. If you prefer a custom role, the essential permission is aiplatform.endpoints.predict, used for both inference and token counting. A dedicated GCP project for Claude Code makes costs and access much easier to reason about.

1M context

Sonnet 5, Opus 4.6 and later, and Sonnet 4.6 support a 1M-token window. Sonnet 5 always uses it, with no [1m] variant. For the others, choose the 1M option in the wizard or append [1m] to a pinned ID. See model configuration.

Troubleshooting

ProblemWhat to try
"Could not load the default credentials"gcloud auth application-default login, or set GOOGLE_APPLICATION_CREDENTIALS to a service account key file
Quota errorsCheck or raise quotas in the Cloud Console
404 "model not found"Confirm the model is enabled in Model Garden and offered at your location (some are only on global, eu or us). With CLOUD_ML_REGION=global, check the model supports global endpoints; if not, pick a different model via ANTHROPIC_MODEL or ANTHROPIC_DEFAULT_HAIKU_MODEL, or route it with its VERTEX_REGION_CLAUDE_* variable
429 errorsOn a regional endpoint, make sure both the primary and small/fast models are offered there; consider CLOUD_ML_REGION=global