Fast mode
Run Opus up to 2.5x faster for a higher per-token price, and understand the billing, requirements, gateway checks and rate limits that come with it.
Fast mode runs Claude Opus on a speed-optimised API configuration. Responses arrive up to 2.5 times faster, and you pay more per token for the privilege. The model, its quality and its capabilities are identical: only the latency and the price change.
I switch it on when I am pairing live with Claude on a bug and every pause costs me focus, and I switch it off before kicking off anything long and unattended.
Note: Fast mode is a research preview. Pricing, availability and behaviour can change.
Which models support it
Fast mode works on Opus 5.5, Opus 5 and Opus 4.8. It does not exist for Sonnet, Haiku, Fable or any other model.
Opus 4.7 lost fast mode support: it was deprecated on 25 June 2026 and removed on 24 July 2026. Switching to Opus 4.7 turns fast mode off.
From v2.1.280, Opus 5.5 is the model fast mode moves you to. Earlier versions used Opus 5 (v2.1.219 onwards), Opus 4.8 (v2.1.154 to v2.1.218) and Opus 4.7 (v2.1.142 to v2.1.153).
Turning it on and off
In the terminal you have two routes:
- Run
/fast, pressSpaceto flip the toggle, thenEnterto confirm. - Put
"fastMode": truein your user settings file (~/.claude/settings.json).
In the VS Code extension, use the Toggle fast mode command, which appears when the selected model supports it. Both routes save to the fastMode setting.
When fast mode comes on:
- If you were on a model that does not support it, Claude Code moves you to Opus.
- You see Fast mode ON.
- A small
↯symbol sits beside the prompt while it is active.
Run /fast at any time to see whether it is on. Turning it off leaves you on Opus; use /model if you want to go back to something else.
You can toggle mid-turn. The turn already in progress finishes at its original speed and the change applies from your next turn. If enabling fast mode also changes your model, the new model is used from the next request within that turn.
By default the setting persists, so a new session starts with fast mode on if you left it on. Organisations can change that with per-session opt-in.
Fast mode in -p runs
In non-interactive mode, /fast only works when the run was launched with fast mode in its --settings JSON (v2.1.205 and later). The toggle then applies to that run and is not saved:
claude -p --settings '{"fastMode": true}' "summarise the failing tests in ci.log"
Anywhere else in -p mode, /fast reports that fast mode is not available.
Fast mode in cloud sessions
Cloud sessions, whether on Anthropic infrastructure or a self-hosted runner, support fast mode from v2.1.271. Type /fast on in the session; it lasts for that session only. In the browser at claude.ai/code the model menu on the message box also shows a fast mode switch when your plan and model allow it. See Claude Code on the web and self-hosted environments.
Switching models with fast mode on
Fast mode follows you as you change models:
- Moving to an unsupported model (including Opus 4.7) turns it off. Before v2.1.221 it stayed on for Opus 4.7 and the API rejected the requests.
- Moving back to a supported Opus turns it back on, but only if your saved preference is on and per-session opt-in is not in force. Otherwise run
/fastagain.
Each of these transitions prints Fast mode ON or Fast mode OFF, whether you switch with /model, /config model=<model>, or from a device connected through Remote Control. Remote Control clients get the updated status after a model switch, a reconnect, or a failed availability check.
What it costs
| Model | Input per million tokens | Output per million tokens |
|---|---|---|
| Opus 5.5 | $8 | $40 |
| Opus 5 | $10 | $50 |
| Opus 4.8 | $10 | $50 |
The rate is flat across the full 1M context window.
The hidden cost is the switch itself. The first time fast mode turns on in a conversation, the whole existing context is charged at the fast mode uncached input rate. Fifty thousand tokens in, that is cheap. Six hundred thousand tokens in, it is not. You pay it once per conversation, so toggling off and on again later does not repeat it. The practical rule: decide at the start. More detail in prompt caching.
On subscription plans (Pro, Max, Team, Enterprise) fast mode is billed only to usage credits. It never comes out of your included allowance, even if you have allowance left.
Finding the spend
Run /status first. A Login method row like Claude Max account means a subscription; an API key row means a Console organisation.
- Pro and Max: the Usage credits section of Settings > Usage on claude.ai. Fast mode is included in the monthly total, not shown separately.
- Team and Enterprise: your organisation pays from its usage credits. Run
/usageto see your own share. See costs for the admin view. - Claude Console: on the Usage and Cost pages, choose Speed (Research Preview) under Group by. The option only appears if the date range contains fast mode usage.
When it is worth it
| Fast mode suits | Standard speed suits |
|---|---|
| Tight edit, run, look loops | Long autonomous tasks |
| Live debugging with someone watching | CI pipelines and batch jobs |
| Deadline work where minutes matter | Anything where cost is the main constraint |
Fast mode versus lower effort
Both make responses quicker, in different ways:
| Lever | Speed | Quality | Cost |
|---|---|---|---|
| Fast mode | Faster | Unchanged | Higher per token |
| Lower effort level | Faster | Can drop on hard problems | Fewer tokens |
They stack. For straightforward edits I sometimes run fast mode at low effort and get near-instant turns.
Requirements
Fast mode needs every one of these:
- Anthropic API or a Claude subscription. It is not offered on Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS.
- Usage credits on subscription plans. Without them,
/fastsays "Fast mode requires usage credits". On Pro and Max, enable them under Settings > Usage on claude.ai or run/usage-credits. On Team and Enterprise, someone with billing access enables them under Organization settings > Usage; anyone else can run/usage-creditsto request it. - A paid Console organisation for API users. On the free Evaluation plan,
/fastreports "Fast mode unavailable during evaluation. Please purchase credits." - Owner enablement on Team and Enterprise, where it is off by default.
Why /fast might refuse
| Message | Cause |
|---|---|
| "Fast mode has been disabled by your organization" | Not enabled for the organisation, or managed settings set fastMode: false, or managed fastModePerSessionOptIn outside an interactive terminal |
| "...is not in your organization's allowed models" | availableModels excludes the Opus model fast mode would switch to. If you are already on an allowed Opus that supports it, /fast enables it on that model instead |
| "Fast mode requires usage credits" | Subscription without usage credits turned on |
Enabling it for an organisation
- Console (API customers): an admin turns it on in Claude Code preferences in the Console. Because it is a research preview, the organisation also needs fast mode access provisioned, via your account manager or the waitlist. Until then every fast request returns a 429, which Claude Code treats as a fast mode rate limit that never clears.
- Team and Enterprise: an Owner enables it under Organization settings > Claude Code on claude.ai.
To remove fast mode completely on a machine, set CLAUDE_CODE_DISABLE_FAST_MODE=1. See environment variables.
Behind proxies and LLM gateways
Before offering fast mode, Claude Code asks api.anthropic.com directly whether your organisation has it. That check ignores ANTHROPIC_BASE_URL but does use an HTTP proxy if one is configured. A previously successful check is cached, so problems mostly hit fresh installs.
| Symptom | Likely cause | Fix |
|---|---|---|
| "Fast mode unavailable due to network connectivity issues" | Direct egress to api.anthropic.com blocked, or the check sent a gateway-issued key (from ANTHROPIC_API_KEY or an apiKeyHelper) that Anthropic rejected | Allowlist api.anthropic.com, or set CLAUDE_CODE_SKIP_FAST_MODE_NETWORK_ERRORS=1 |
"Fast mode has been disabled by your organization" with ANTHROPIC_AUTH_TOKEN only | No claude.ai login or Anthropic key, so the check is skipped and treated as disabled | CLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK=1 |
| Same message behind a TLS-inspecting proxy | The proxy answered with its own 200 page | CLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK=1 |
| "Fast mode is currently unavailable" | CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC suppressed the check | Either skip variable |
CLAUDE_CODE_SKIP_FAST_MODE_NETWORK_ERRORS treats a failed check as success but still respects a genuine "disabled" answer. CLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK skips the check outright. Neither overrides the API: if your organisation really has fast mode off, the API rejects the request, Claude Code retries it at standard speed, and fast mode switches off. See network configuration and LLM gateways.
Require per-session opt-in
Persistent fast mode is convenient for one session and expensive across six parallel ones. Set fastModePerSessionOptIn in any settings file to make every session start with it off:
{
"fastMode": false,
"fastModePerSessionOptIn": true
}
The user's saved preference is kept, so removing the key restores the old behaviour. Team and Enterprise Owners can push it to everyone through server-managed settings.
When it comes from managed settings, /fast on only works in an interactive terminal. In -p mode, VS Code and cloud sessions it is refused with the "disabled by your organization" message.
Rate limits and running out of credits
Fast mode has its own rate limit pool, shared by every supported Opus model. When you hit it:
- Requests drop back to standard speed and price.
- The
↯symbol turns grey while the cooldown runs. - Fast mode comes back automatically when the cooldown ends.
Run /fast to switch it off yourself instead of waiting.
Running out of usage credits mid-session is different: each rejected fast request is retried at standard speed, with no cooldown. Interactive sessions show Fast mode disabled · usage credits exhausted and switch fast mode off for the session without touching your saved preference. In -p mode with --output-format stream-json, and through the Agent SDK, the same text arrives once per turn as a system message with subtype notification, and fast mode stays on (v2.1.221+).