Voice Input for Agentic Coding
Voice input for agentic coding means speaking prompts instead of typing them. Claude Code transcribes speech into the prompt with /voice (claude.ai login required), Codex CLI runs a spoken voice conversation, GitHub Copilot CLI toggles dictation with Ctrl+Space, and system dictation covers every other tool. The transcript still needs checking before it becomes a decision.
This page is for the developer who spends more time explaining intent to an agent than reviewing its code, and the tech lead who has to decide whether microphones belong in the team’s setup. You are 40 minutes into a /grill-with-docs session about per-seat billing. Question 17 asks what happens to a seat removed mid-cycle, and the honest answer has three conditions and an exception. You type “prorate it”. The agent fills in the rest with its own recommendation. Three days later the refund logic is wrong in exactly the way you would have explained if typing were cheaper.
What you get from voice input set up properly
Section titled “What you get from voice input set up properly”- A decision table of voice options per agent, and what each one sends where.
- A working Claude Code setup:
/voice tap, a persisted setting, the right dictation language, and a rebound key. - A framework-level workflow: answer a grill-me interview by voice, then turn the spoken answers into a readback, acceptance criteria, and failing tests you can check without replaying the audio.
- Three copy-paste prompts that make an agent cope with transcription errors instead of guessing around them.
- A team checklist, the traps, and the failure modes with fixes.
Which voice input should you use with each agent?
Section titled “Which voice input should you use with each agent?”Start from the agent you already use. The built-in options differ in kind: Claude Code and Copilot CLI turn speech into prompt text, while Codex CLI holds a spoken conversation.
| Agent | Built-in voice | What it does | Needs | Checked against |
|---|---|---|---|---|
| Claude Code (CLI, agent view, VS Code extension) | /voice, /voice hold, /voice tap, /voice off | Live transcription into the prompt; you can mix typing and speech | claude.ai login, local microphone, WSLg if you run in WSL | Claude Code voice dictation docs, 2026-09-26 |
| Codex CLI | /voice, /voice settings, /voice mute, /voice stop, F8 | Spoken conversation; Codex listens and answers aloud | macOS, an MSVC-based Windows build, or a glibc-based Linux build | codex-cli 0.157.1 binary and release notes 0.155.0 and 0.156.0 |
| GitHub Copilot CLI | Ctrl+Space, /voice devices | Toggles voice dictation; /voice devices picks and remembers the microphone | A voice runtime the CLI downloads | github/copilot-cli changelog 1.0.71 and 1.0.81 |
| Cursor | Voice input in Agent (secondary sources only, see the Cursor tab) | Speech into the Agent prompt | Not verified for this page | Not verified: cursor.com unreachable on 2026-09-26 |
| Aider | /voice, then Enter | Records, transcribes, and sends; mature, with slowing releases (last PyPI release 2026-02-12) | Aider’s optional voice dependencies | Aider usage/voice.md and PyPI aider-chat, 2026-09-26 |
| Any tool | macOS Dictation, Windows voice typing (Win+H) | Types into whatever window has focus | OS support | OS vendor |
Every built-in option in the table streams audio to a vendor. Claude Code’s documentation says: “Audio is not processed locally.” Copilot CLI’s changelog does not say where transcription runs, so treat it as cloud until you confirm otherwise.
Popularity as of 2026-09-26: no vendor publishes usage numbers for voice input, and there is no install count because the feature ships inside each agent (this site’s agent-tools dossier found no public metric). What can be dated is release activity: Claude Code’s changelog lists 20 dictation languages from 2.1.69 (published to npm on 2026-03-04), Copilot CLI added Ctrl+Space dictation in 1.0.81 (2026-08-27), and Codex 0.156.0 (2026-09-22) turned voice conversations on by default.
Set up Claude Code voice dictation in tap mode
Section titled “Set up Claude Code voice dictation in tap mode”Tap mode suits long answers: you tap once, speak for as long as you need, and tap again to send. Unlike hold mode, it needs no key-repeat warmup, which fails in terminals with key repeat off.
-
Confirm you are signed in with a claude.ai account. Run
/statusin Claude Code. If the login method is an API key, Bedrock, Google Cloud’s Agent Platform, or Microsoft Foundry, dictation is not available; run/loginand sign in with claude.ai. -
Turn on dictation in tap mode:
/voice tapWhen you enable it, Claude Code runs a microphone check. On macOS the first run triggers the system permission prompt for your terminal app.
-
With the prompt empty, tap
Space. The footer shows● REC · tap to send. Speak, then tapSpaceagain. Claude Code inserts the transcript and submits it when it is at least three words long, so a stray tap never sends one word. Recording also stops after 15 seconds of silence or two minutes in total.EscorCtrl+Cdiscards the recording and restores the prompt. -
To dictate in Polish (or another of the 20 supported languages), set
languagein/configor in~/.claude/settings.json. The same setting makes Claude answer in that language and name sessions in it:{"language": "polish","voice": {"enabled": true,"mode": "tap"}}/voicewrites thevoiceobject for you, and it persists across sessions. In hold mode, add"autoSubmit": trueto send on key release; without it, you pressEnteryourself. -
Optional: move dictation off
Space. Bindvoice:pushToTalkto another key in~/.claude/keybindings.json; the new key replacesSpace(the"space": nullline only makes that explicit). A modifier combination such asmeta+kstarts recording on the first keypress in both modes, with no warmup:{"bindings": [{"context": "Chat","bindings": {"meta+k": "voice:pushToTalk","space": null}}]}
You should see your words appear dimmed in the prompt as you speak, then turn solid when the transcript is final. Claude Code adds your project name and git branch name as recognition hints and is tuned for terms such as OAuth, JSON, and localhost. Your own table and function names are not on that list, which is why the workflow below checks them.
Transcription does not consume Claude messages or tokens and does not count toward the limits in /usage, according to Claude Code’s documentation. Context cost is the same as for a typed prompt of the same length, and a two-minute spoken answer can add a few hundred words.
Voice input in Codex, Cursor, Copilot CLI, and other tools
Section titled “Voice input in Codex, Cursor, Copilot CLI, and other tools”Use the setup above. Dictation also works in agent view: hold or tap your push-to-talk key while the dispatch input or a peek-panel reply is focused to dictate to a background session. The VS Code extension supports dictation with the same claude.ai requirement, but not in VS Code Remote sessions (SSH, Dev Containers, and Codespaces), because the microphone is on your machine and the extension runs on the remote host.
In codex-cli 0.157.1, /voice in the TUI starts a spoken voice conversation rather than dictation into the prompt: Codex listens and answers aloud. The in_app_dictation feature flag is on, but codex-cli 0.157.1 exposes no TUI command for it (checked 2026-09-26). Release 0.156.0 turned voice conversations on by default and added the F8 toggle.
| Command | What it does |
|---|---|
/voice or F8 | “Start or stop a voice conversation.” (binary wording) |
/voice settings | Chooses a voice; the binary says the choice “Applies to your next voice conversation.” |
/voice mute | “Mute or unmute the voice microphone” (binary wording) |
/voice stop | Ends the voice conversation (our description; the binary lists the form without one) |
The four forms come from the binary’s usage line, Usage: /voice [settings|mute|stop]. Its messages say where voice will not start: “Start a conversation before using voice mode.”, “Voice mode is unavailable in side conversations.”, and it “requires macOS, an MSVC-based Windows build, or a glibc-based Linux build”. codex-cli 0.157.1 from npm on Linux x64 is a musl build, so voice is off there (checked 2026-09-26); use macOS, Windows, or a glibc-based Codex build.
codex features list shows realtime_conversation and in_app_dictation, both stable and on in 0.157.1.
Cursor’s changelog, as quoted by secondary sources, describes voice input for Agent since Cursor 2.0 and “press and hold Ctrl+M” since Cursor 3.1. We could not confirm either against cursor.com, which was unreachable from our environment on 2026-09-26. If Cursor’s keyboard shortcuts show no voice binding, use system dictation in the Agent input.
GitHub Copilot CLI toggles voice dictation with Ctrl+Space (changelog 1.0.81, 2026-08-27). Pick and persist the microphone with:
/voice devicesThe first use downloads a voice runtime. Since 1.0.76, voice mode pauses playing media while it records on macOS and Windows.
System dictation types into whatever window has focus, so it works in an agent that has no voice of its own, in an SSH session, and in Claude Code on an API key, Bedrock, or a gateway or local model, where /voice is unavailable. Turn on macOS Dictation under System Settings › Keyboard › Dictation; on Windows, press Win+H for voice typing.
What you lose: no project-name or branch hints, and no protection against sending a stray word. Commercial dictation apps add code-aware vocabulary but are another processor of your speech; clear them with security like any new SaaS tool.
Dictate your grill-me answers: the workflow
Section titled “Dictate your grill-me answers: the workflow”A /grill-me or /grill-with-docs session is the best fit for voice because the agent asks a round of numbered questions, each with a recommended answer, and waits for yours. Your job is to accept, reject, or qualify each recommendation, and the qualification is the part people drop when they have to type it. The risk is that the transcript, not your intent, becomes the spec. The workflow keeps the typing for the parts that must be exact and puts a check between speech and code.
Install the interview skills in your agent
Section titled “Install the interview skills in your agent”grill-with-docs ships in Matt Pocock’s mattpocock/skills pack. Install it once, with the plugin or the skills CLI, not both: the pack’s README warns that every skill then appears twice. The invocation differs per tool.
# terminal: official marketplace, nothing to add firstclaude plugin install mattpocock-skillsPlugin skills are namespaced, so you start the interview with /mattpocock-skills:grill-with-docs. Run /mattpocock-skills:setup-matt-pocock-skills once per repository first; it asks where the glossary and ADRs go. The plugin (1.2.3, 25 skills) adds about 1,600 tokens of always-loaded skill descriptions, measured by this site’s frameworks dossier on 2026-09-26; check yours with /context.
# terminal, in the repository rootnpx skills add mattpocock/skills --skill grill-with-docs grilling domain-modeling setup-matt-pocock-skills -a codexCodex invokes skills with $: run $setup-matt-pocock-skills once, then start the interview with $grill-with-docs.
# terminal, in the repository rootnpx skills add mattpocock/skills --skill grill-with-docs grilling domain-modeling setup-matt-pocock-skills -a cursorIn the Agent input, run /setup-matt-pocock-skills once, then start the interview with /grill-with-docs.
Run the interview by voice
Section titled “Run the interview by voice”-
Start the session by typing. Type your tool’s command from the tabs above, a space, and then paste the first prompt below. Commands, file paths, and identifiers are faster and safer typed than spoken.
-
Answer each question by voice. In Claude Code with
/voice tap, tap, speak and tap. The agent asks a numbered round, so start each spoken answer with its number, then use the same four-part shape: the decision, the reason, the constraint, and what would change your mind. “Q3: prorate to the day, because finance reconciles daily, except annual plans, which keep the seat until renewal; I would change this if Stripe cannot do per-day credits.” -
Type the identifiers. In tap mode the second tap sends the answer, so it cannot pause for typing. When an answer must name a table, column, function, or flag, switch to
/voice hold(withoutautoSubmit): holdSpaceand speak, release, type the name, hold again to continue, and pressEnterwhen the answer is complete. Claude Code inserts each transcript at the cursor, so speech and typing mix in one prompt. -
Ask for a readback before anything is written. At the end of the interview, use the second prompt below. The agent lists every decision with the words you actually said and flags anything that looks misheard. Read that list, not the transcript.
-
Turn decisions into things a machine checks. Use the third prompt to write acceptance criteria and failing tests. With
/grill-with-docs, the resolved terms land inGLOSSARY.mdand hard decisions in ADRs; review those diffs withgit diff, because they outlive the conversation. Repositories set up with an older version of the skills may still haveCONTEXT.md, the name the skills used before they switched toGLOSSARY.md(mattpocock/skills README, checked 2026-09-26); rename it withgit mv. -
Sign off on the artifacts, not the audio. You approve the readback, the acceptance criteria, and the red tests. From here the build runs like any spec-driven change: the tests going green is the proof, and nobody replays what you said.
How do you know the agent heard what you meant?
Section titled “How do you know the agent heard what you meant?”You do not check the transcript word by word. You check three artifacts that a misheard word cannot slip through:
- The readback list. Every decision in one sentence beside your quoted words. A wrong quote is a transcription error; a right quote with a wrong decision is a misunderstanding. Both surface here, before any code exists.
- The acceptance criteria and red tests. A misheard “annual” versus “and you’ll” becomes a Given/When/Then line you can read in seconds, and a test that fails for the stated reason proves the criterion is executable.
- The
GLOSSARY.mdand ADR diffs from/grill-with-docs. Terms in the glossary must match the names in the code. A term that exists nowhere in the repository is almost always a transcription artifact.
The developer who ran the interview signs off on the readback and the red tests. On a team, the reviewer of the pull request reads the acceptance criteria and the ADRs, not the conversation log.
Roll out voice input to a team
Section titled “Roll out voice input to a team”For a tech lead, the questions are data handling and where it works, not whether people like it.
| Check | Why it matters | What to do |
|---|---|---|
| Where does audio go? | Claude Code streams audio to Anthropic for transcription; Codex voice is a live conversation with OpenAI’s service | Treat speech like prompt text under your AI data policy; read Claude Code’s data-usage page |
| Which sign-in does your team use? | Claude Code dictation needs a claude.ai login; API-key, Bedrock, Agent Platform, and Foundry setups cannot use it | If you standardised on a cloud provider, plan on system dictation instead |
| Where do people work? | No dictation over SSH, in cloud sessions, or in VS Code Remote, Dev Containers, or Codespaces | Remote-first setups use system dictation on the local machine |
| Can it be turned off? | An administrator policy can disable Claude Code voice dictation; users then see Voice mode is disabled by your organization's policy | Decide before rollout, not after the first complaint |
| Shared offices | Spoken prompts include customer names and incident details | Headsets and push-to-talk (hold mode) in open-plan spaces |
What breaks with voice dictation, and how to fix it
Section titled “What breaks with voice dictation, and how to fix it”| Symptom | Cause | Fix |
|---|---|---|
Voice mode requires a Claude.ai account | Signed in with an API key or a third-party provider | /login with claude.ai, or use system dictation |
Holding Space types spaces and nothing records | Dictation is off, or the terminal sends no key-repeat events | /voice hold to turn it on; if one or two spaces appear and stop, switch to /voice tap |
Tapping Space types a space in tap mode | The first tap only records when the prompt is empty | Clear the prompt, or rebind to meta+k |
Microphone access is denied | The terminal app lacks microphone permission | macOS: System Settings › Privacy & Security › Microphone; Windows: allow microphone access for desktop apps; run /voice again |
| Terminal missing from the macOS microphone list | Stale permission state | tccutil reset Microphone com.apple.Terminal (or your terminal’s bundle ID), then quit (Cmd+Q) and relaunch the terminal |
Voice mode requires SoX for audio recording on Linux | Native audio module did not load and no fallback exists | Install SoX with the command the error shows |
could not find a working audio recorder in WSL | SoX lacks its PulseAudio backend | sudo apt install sox libsox-fmt-pulse |
Voice input is failing repeatedly and has been paused | Three failures in 10 seconds, usually no capture device (headless host, remote shell) | Fix the device, wait 10 seconds, try again |
Voice stream error: WebSocket upgrade rejected with HTTP 4xx | Stale sign-in, or a proxy or bot-protection service answering instead | /login; check the VPN or proxy on the path |
| Garbled text or the wrong language | Dictation defaults to English | Set language in /config first |
Codex /voice refuses to start | No conversation yet, a side conversation, or an unsupported build | Send a first message or return to the main thread. Unsupported build: codex-cli 0.157.1 on Linux x64 from npm is musl (see the Codex tab) |
If a misheard answer already reached the agent, do not argue with it in the next turn. In Claude Code, press Esc twice with an empty prompt to open the rewind menu and restore the conversation before that answer, or type the correction explicitly (“Correction to decision 4: the column is seat_count, not seed count”) and ask for the readback again.
Where to go next with voice-driven planning
Section titled “Where to go next with voice-driven planning”- Grill Me and Grill With Docs for the interview skills this workflow drives.
- Acceptance criteria that agents can verify to turn the confirmed decisions into checks.
- Spec-driven development for the build loop after the tests are red.
- Plan mode to gate execution once the interview is done.
- Remote and mobile agent clients for steering agents away from your desk.
- The agent tools overview for the rest of the tooling around the agent.
Frequently asked questions
How do I dictate prompts in Claude Code?
Run /voice to turn dictation on, then hold Space while you speak, or run /voice tap to tap Space once to start and again to send. You need a claude.ai login and a local microphone; it does not work with an API key, Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, SSH, or cloud sessions.
Can I dictate to Claude Code in Polish?
Yes. Dictation supports 20 languages including Polish. It follows the language setting, which also sets the language Claude answers in, so setting it to Polish changes both.
Does Codex CLI have voice dictation?
Not as prompt dictation. In codex-cli 0.157.1, /voice in the TUI starts a spoken voice conversation rather than dictation into the prompt; the in_app_dictation feature flag is on, but codex-cli 0.157.1 exposes no TUI command for it (checked 2026-09-26). /voice starts or stops the conversation, /voice settings picks a voice, and F8 toggles it. It requires macOS, an MSVC-based Windows build, or a glibc-based Linux build.
Does voice dictation cost tokens?
Not in Claude Code: its documentation says transcription does not consume Claude messages or tokens and does not count toward the /usage limits. The prompt the transcript becomes is billed like any typed prompt.