Voice dictation
Speak prompts into the Claude Code CLI with hold-to-talk or tap-to-send dictation, change the language, rebind the key and fix microphone problems.
Sometimes it is quicker to say what you want than type it, especially for long, rambling task descriptions. Claude Code's voice dictation transcribes your speech straight into the prompt as you talk, so you can mix spoken and typed text in one message. Switch it on with /voice, then either hold a key while speaking or tap once to start and once to send.
Hold mode also works in agent view: hold your push-to-talk key while the dispatch box or a peek-panel reply is focused to dictate to a background session.
What you need
Audio is streamed to Anthropic's servers for transcription; nothing is transcribed locally. That brings three requirements:
- A claude.ai sign-in. Dictation is unavailable with a direct Anthropic API key, Amazon Bedrock, Google Vertex AI or Microsoft Foundry.
- A microphone on the machine running Claude Code. It will not work in cloud sessions or over SSH.
- WSLg if you are on WSL. WSLg ships with WSL2 from the Microsoft Store on Windows 10 and 11. On WSL1 or without WSLg, run Claude Code natively on Windows.
Transcription does not use Claude messages or tokens and does not count against the limits in /usage. For data handling see data usage.
Recording uses a built-in native module on macOS, Linux and Windows. On Linux, if that module fails to load, Claude Code falls back to arecord (ALSA utils) or rec (SoX), and /voice prints an install command if neither is present.
The VS Code extension supports dictation with the same claude.ai requirement, but not in Remote sessions (SSH, Dev Containers, Codespaces), because the microphone is local and the extension runs remotely.
Switching it on
Run /voice. It checks the microphone, which on macOS triggers the system permission prompt the first time. You will see a confirmation along these lines:
Voice mode enabled (hold). Hold space to record. Dictation language: en (/config to change).
| Command | What it does |
|---|---|
/voice | Toggle on or off, keeping the current mode |
/voice hold | Turn on in hold mode |
/voice tap | Turn on in tap mode |
/voice off | Turn off |
The choice persists across sessions. You can also set it in your user settings:
{
"voice": {
"enabled": true,
"mode": "hold",
"autoSubmit": true
}
}
For your first three sessions with voice on, an empty prompt shows a hold space to speak hint in the footer (it reflects your actual voice:pushToTalk key, reads the same in both modes, and is hidden if you use a custom status line).
Recognition is tuned for developer vocabulary, so terms such as regex, OAuth, JSON and localhost come through properly. Your project name and current git branch are fed in as hints automatically.
Hold mode
Hold mode is push-to-talk and is the default. Hold Space, talk, let go.
Terminals do not report "key held", so Claude Code infers it from the burst of key-repeat events. That means a short warmup: the footer shows keep holding…, then listening… once recording starts. The first couple of repeated spaces land in the input during warmup and are deleted automatically. A quick single tap still types a normal space. While recording, the cursor becomes a level meter that bobs with your voice, unless prefersReducedMotion is on.
Words appear dimmed while you speak and firm up when you release. The text is inserted at the cursor and the cursor ends up after it, so you can type, dictate, move the cursor, dictate again:
> add a migration that ▮
(hold space: "adds a nullable archived_at column to the projects table")
> add a migration that adds a nullable archived_at column to the projects table▮
By default you still press Enter to send. With "autoSubmit": true in the voice object, releasing the key sends the prompt as long as the transcript is three words or more.
Space only starts dictation where it would otherwise type into the prompt. In the transcript viewer it pages, and in Vim mode outside INSERT it is a command. A modifier chord such as meta+k never types anything, so it works from those places too.
Tip: Dislike the warmup? Use tap mode, or bind a modifier chord such as
meta+k, which starts recording on the first press.
Tap mode
Tap once to start, speak, tap again to stop and send. No warmup and no holding. Enable it with /voice tap.
With the prompt empty, tap Space; the footer shows ● REC · tap to send. Tap again to finish. If the transcript is three words or longer it is submitted automatically; anything shorter is inserted but not sent, so a stray tap cannot fire off "um". For Japanese, Chinese and Thai, which are written without spaces, the three-word rule counts actual words, so auto-submit works there too (in tap mode and in hold mode with autoSubmit).
The first tap only starts recording on an empty prompt, so you can keep typing spaces normally mid-message. The second tap stops regardless. Recording also stops itself after 15 seconds of silence or two minutes in total.
Cancelling
Esc or Ctrl+C abandons a recording: the microphone stops, the transcript is thrown away and the prompt returns to what it held before. Both also work while a finished recording is still being processed; if you already edited or sent the prompt by then, your edit stands.
That cancelling keypress does nothing else. Esc will not interrupt Claude, and Ctrl+C will not clear the prompt or count towards the double press that quits.
Dictation language
Dictation follows the language setting, the same one that sets Claude's reply language. Empty means English. In the VS Code extension an empty language falls back to VS Code's accessibility.voice.speechLanguage before English.
Set it in /config or in settings, using either a BCP 47 code or the language name:
{
"language": "de"
}
Supported languages and codes:
| Language | Code | Language | Code |
|---|---|---|---|
| Czech | cs | Japanese | ja |
| Danish | da | Korean | ko |
| Dutch | nl | Norwegian | no |
| English | en | Polish | pl |
| French | fr | Portuguese | pt |
| German | de | Russian | ru |
| Greek | el | Spanish | es |
| Hindi | hi | Swedish | sv |
| Indonesian | id | Turkish | tr |
| Italian | it | Ukrainian | uk |
If your language is not on the list, /voice warns you and dictation falls back to English. Claude's written replies still use your chosen language.
Changing the dictation key
The key is the voice:pushToTalk action in the Chat context, Space by default, and it drives both modes. Rebind it in ~/.claude/keybindings.json (see keybindings):
{
"bindings": [
{
"context": "Chat",
"bindings": {
"ctrl+shift+v": "voice:pushToTalk"
}
}
]
}
Only one key triggers voice:pushToTalk at a time, so a custom binding replaces Space automatically; you do not need to null it out.
In hold mode, avoid a bare letter like d: hold detection needs key-repeat, so the letter would type into the prompt during warmup. Stick with Space or a modifier chord. Tap mode has no warmup, so most keys are fine. Keys the terminal never receives, such as Caps Lock, cannot be bound and produce an error.
Troubleshooting
| Message or symptom | Cause and fix |
|---|---|
Unknown command: /voice | /voice only exists while a claude.ai account is the active sign-in. Run /login. If ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, an apiKeyHelper or a third-party provider is configured, it takes priority; remove it and restart. |
Voice mode requires a Claude.ai account | No usable claude.ai sign-in was found. Run /login. |
Voice mode is disabled by your organization's policy | An admin has turned it off. Ask your administrator. |
Microphone access is denied | Grant the terminal microphone access. macOS: System Settings → Privacy & Security → Microphone. Windows: Settings → Privacy & security → Microphone, allow desktop apps. Then run /voice again. |
Voice mode requires SoX for audio recording (Linux) | Native module failed and no fallback exists. Install SoX, for example sudo apt-get install sox. |
Voice mode requires a microphone, but SoX could not open an audio capture device | SoX is present but the host has no capture device (headless server, container). Use a machine with a microphone. Reported this way from v2.1.195. |
Voice mode could not find a working audio recorder in WSL | WSLg routes audio via PulseAudio. Run sudo apt install sox libsox-fmt-pulse; plain sox only brings the ALSA backend, and WSL has no /dev/snd. |
Voice input is failing repeatedly and has been paused | Three failures within 10 seconds pauses dictation until 10 seconds after the first. Usually no working audio input. Fix the root cause and try again. (Before v2.1.202 only start-up failures counted.) |
Holding Space does nothing in hold mode | Watch the prompt. Spaces piling up means voice is off: run /voice hold. One or two spaces then nothing means key-repeat is not reaching Claude Code (perhaps disabled in the OS); use /voice tap. |
Tapping Space types a space in tap mode | The first tap only records on an empty prompt. Clear it, or confirm the mode with /voice tap. |
No audio detected from microphone | Silence captured. Check the default input device and its level (Windows: Settings → System → Sound → Input; macOS: System Settings → Sound → Input). |
Voice connection failed | The audio never reached the service. Check your network. (Before v2.1.200 a silent mic could wrongly show this.) |
Voice stream error: WebSocket upgrade rejected with HTTP <status> | A server refused the connection. A 4xx usually means a stale sign-in or a proxy/bot-protection layer answering instead; run /login and check VPNs and proxies. Non-4xx statuses are retried once if you are still recording. In v2.1.229 to v2.1.231 native builds showed Voice connection failed instead. |
No speech detected | Audio arrived but no words were recognised. Get closer to the mic, cut background noise, check the dictation language. |
| Garbled or wrong-language text | Dictation defaults to English. Set language first. |
Terminal missing from macOS Microphone settings
If your terminal is not listed, there is nothing to toggle. Reset its permission so macOS asks afresh:
- Run
tccutil reset Microphone <bundle-id>. Usecom.apple.Terminalfor Terminal orcom.googlecode.iterm2for iTerm2; find others withosascript -e 'id of app "Ghostty"'(swap in your app's name). - Quit the terminal completely with
Cmd+Q(closing windows is not enough, macOS will not re-prompt a running process) and reopen it. - Start Claude Code, run
/voiceand allow access when prompted.
Warning:
tccutil reset Microphonewith no bundle ID revokes microphone access for every app on the Mac, Zoom and Slack included. Do not run it mid-call.