Skip to content

Fast mode

Run Opus up to 2.5x faster for a higher per-token price, and understand the billing, requirements, gateway checks and rate limits that come with it.

Fast mode runs Claude Opus on a speed-optimised API configuration. Responses arrive up to 2.5 times faster, and you pay more per token for the privilege. The model, its quality and its capabilities are identical: only the latency and the price change.

I switch it on when I am pairing live with Claude on a bug and every pause costs me focus, and I switch it off before kicking off anything long and unattended.

Note: Fast mode is a research preview. Pricing, availability and behaviour can change.

Which models support it

Fast mode works on Opus 5.5, Opus 5 and Opus 4.8. It does not exist for Sonnet, Haiku, Fable or any other model.

Opus 4.7 lost fast mode support: it was deprecated on 25 June 2026 and removed on 24 July 2026. Switching to Opus 4.7 turns fast mode off.

From v2.1.280, Opus 5.5 is the model fast mode moves you to. Earlier versions used Opus 5 (v2.1.219 onwards), Opus 4.8 (v2.1.154 to v2.1.218) and Opus 4.7 (v2.1.142 to v2.1.153).

Turning it on and off

In the terminal you have two routes:

  • Run /fast, press Space to flip the toggle, then Enter to confirm.
  • Put "fastMode": true in your user settings file (~/.claude/settings.json).

In the VS Code extension, use the Toggle fast mode command, which appears when the selected model supports it. Both routes save to the fastMode setting.

When fast mode comes on:

  • If you were on a model that does not support it, Claude Code moves you to Opus.
  • You see Fast mode ON.
  • A small ↯ symbol sits beside the prompt while it is active.

Run /fast at any time to see whether it is on. Turning it off leaves you on Opus; use /model if you want to go back to something else.

You can toggle mid-turn. The turn already in progress finishes at its original speed and the change applies from your next turn. If enabling fast mode also changes your model, the new model is used from the next request within that turn.

By default the setting persists, so a new session starts with fast mode on if you left it on. Organisations can change that with per-session opt-in.

Fast mode in -p runs

In non-interactive mode, /fast only works when the run was launched with fast mode in its --settings JSON (v2.1.205 and later). The toggle then applies to that run and is not saved:

claude -p --settings '{"fastMode": true}' "summarise the failing tests in ci.log"

Anywhere else in -p mode, /fast reports that fast mode is not available.

Fast mode in cloud sessions

Cloud sessions, whether on Anthropic infrastructure or a self-hosted runner, support fast mode from v2.1.271. Type /fast on in the session; it lasts for that session only. In the browser at claude.ai/code the model menu on the message box also shows a fast mode switch when your plan and model allow it. See Claude Code on the web and self-hosted environments.

Switching models with fast mode on

Fast mode follows you as you change models:

  • Moving to an unsupported model (including Opus 4.7) turns it off. Before v2.1.221 it stayed on for Opus 4.7 and the API rejected the requests.
  • Moving back to a supported Opus turns it back on, but only if your saved preference is on and per-session opt-in is not in force. Otherwise run /fast again.

Each of these transitions prints Fast mode ON or Fast mode OFF, whether you switch with /model, /config model=<model>, or from a device connected through Remote Control. Remote Control clients get the updated status after a model switch, a reconnect, or a failed availability check.

What it costs

ModelInput per million tokensOutput per million tokens
Opus 5.5$8$40
Opus 5$10$50
Opus 4.8$10$50

The rate is flat across the full 1M context window.

The hidden cost is the switch itself. The first time fast mode turns on in a conversation, the whole existing context is charged at the fast mode uncached input rate. Fifty thousand tokens in, that is cheap. Six hundred thousand tokens in, it is not. You pay it once per conversation, so toggling off and on again later does not repeat it. The practical rule: decide at the start. More detail in prompt caching.

On subscription plans (Pro, Max, Team, Enterprise) fast mode is billed only to usage credits. It never comes out of your included allowance, even if you have allowance left.

Finding the spend

Run /status first. A Login method row like Claude Max account means a subscription; an API key row means a Console organisation.

  • Pro and Max: the Usage credits section of Settings > Usage on claude.ai. Fast mode is included in the monthly total, not shown separately.
  • Team and Enterprise: your organisation pays from its usage credits. Run /usage to see your own share. See costs for the admin view.
  • Claude Console: on the Usage and Cost pages, choose Speed (Research Preview) under Group by. The option only appears if the date range contains fast mode usage.

When it is worth it

Fast mode suitsStandard speed suits
Tight edit, run, look loopsLong autonomous tasks
Live debugging with someone watchingCI pipelines and batch jobs
Deadline work where minutes matterAnything where cost is the main constraint

Fast mode versus lower effort

Both make responses quicker, in different ways:

LeverSpeedQualityCost
Fast modeFasterUnchangedHigher per token
Lower effort levelFasterCan drop on hard problemsFewer tokens

They stack. For straightforward edits I sometimes run fast mode at low effort and get near-instant turns.

Requirements

Fast mode needs every one of these:

  • Anthropic API or a Claude subscription. It is not offered on Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS.
  • Usage credits on subscription plans. Without them, /fast says "Fast mode requires usage credits". On Pro and Max, enable them under Settings > Usage on claude.ai or run /usage-credits. On Team and Enterprise, someone with billing access enables them under Organization settings > Usage; anyone else can run /usage-credits to request it.
  • A paid Console organisation for API users. On the free Evaluation plan, /fast reports "Fast mode unavailable during evaluation. Please purchase credits."
  • Owner enablement on Team and Enterprise, where it is off by default.

Why /fast might refuse

MessageCause
"Fast mode has been disabled by your organization"Not enabled for the organisation, or managed settings set fastMode: false, or managed fastModePerSessionOptIn outside an interactive terminal
"...is not in your organization's allowed models"availableModels excludes the Opus model fast mode would switch to. If you are already on an allowed Opus that supports it, /fast enables it on that model instead
"Fast mode requires usage credits"Subscription without usage credits turned on

Enabling it for an organisation

  • Console (API customers): an admin turns it on in Claude Code preferences in the Console. Because it is a research preview, the organisation also needs fast mode access provisioned, via your account manager or the waitlist. Until then every fast request returns a 429, which Claude Code treats as a fast mode rate limit that never clears.
  • Team and Enterprise: an Owner enables it under Organization settings > Claude Code on claude.ai.

To remove fast mode completely on a machine, set CLAUDE_CODE_DISABLE_FAST_MODE=1. See environment variables.

Behind proxies and LLM gateways

Before offering fast mode, Claude Code asks api.anthropic.com directly whether your organisation has it. That check ignores ANTHROPIC_BASE_URL but does use an HTTP proxy if one is configured. A previously successful check is cached, so problems mostly hit fresh installs.

SymptomLikely causeFix
"Fast mode unavailable due to network connectivity issues"Direct egress to api.anthropic.com blocked, or the check sent a gateway-issued key (from ANTHROPIC_API_KEY or an apiKeyHelper) that Anthropic rejectedAllowlist api.anthropic.com, or set CLAUDE_CODE_SKIP_FAST_MODE_NETWORK_ERRORS=1
"Fast mode has been disabled by your organization" with ANTHROPIC_AUTH_TOKEN onlyNo claude.ai login or Anthropic key, so the check is skipped and treated as disabledCLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK=1
Same message behind a TLS-inspecting proxyThe proxy answered with its own 200 pageCLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK=1
"Fast mode is currently unavailable"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC suppressed the checkEither skip variable

CLAUDE_CODE_SKIP_FAST_MODE_NETWORK_ERRORS treats a failed check as success but still respects a genuine "disabled" answer. CLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK skips the check outright. Neither overrides the API: if your organisation really has fast mode off, the API rejects the request, Claude Code retries it at standard speed, and fast mode switches off. See network configuration and LLM gateways.

Require per-session opt-in

Persistent fast mode is convenient for one session and expensive across six parallel ones. Set fastModePerSessionOptIn in any settings file to make every session start with it off:

{
  "fastMode": false,
  "fastModePerSessionOptIn": true
}

The user's saved preference is kept, so removing the key restores the old behaviour. Team and Enterprise Owners can push it to everyone through server-managed settings.

When it comes from managed settings, /fast on only works in an interactive terminal. In -p mode, VS Code and cloud sessions it is refused with the "disabled by your organization" message.

Rate limits and running out of credits

Fast mode has its own rate limit pool, shared by every supported Opus model. When you hit it:

  1. Requests drop back to standard speed and price.
  2. The ↯ symbol turns grey while the cooldown runs.
  3. Fast mode comes back automatically when the cooldown ends.

Run /fast to switch it off yourself instead of waiting.

Running out of usage credits mid-session is different: each rejected fast request is retried at standard speed, with no cooldown. Interactive sessions show Fast mode disabled · usage credits exhausted and switch fast mode off for the session without touching your saved preference. In -p mode with --output-format stream-json, and through the Agent SDK, the same text arrives once per turn as a system message with subtype notification, and fast mode stays on (v2.1.221+).