Skip to content

Voice Input for Agentic Coding

Voice input for agentic coding means speaking prompts instead of typing them. Claude Code transcribes speech into the prompt with /voice (claude.ai login required), Codex CLI runs a spoken voice conversation, GitHub Copilot CLI toggles dictation with Ctrl+Space, and system dictation covers every other tool. The transcript still needs checking before it becomes a decision.

This page is for the developer who spends more time explaining intent to an agent than reviewing its code, and the tech lead who has to decide whether microphones belong in the team’s setup. You are 40 minutes into a /grill-with-docs session about per-seat billing. Question 17 asks what happens to a seat removed mid-cycle, and the honest answer has three conditions and an exception. You type “prorate it”. The agent fills in the rest with its own recommendation. Three days later the refund logic is wrong in exactly the way you would have explained if typing were cheaper.

What you get from voice input set up properly

Section titled “What you get from voice input set up properly”
  • A decision table of voice options per agent, and what each one sends where.
  • A working Claude Code setup: /voice tap, a persisted setting, the right dictation language, and a rebound key.
  • A framework-level workflow: answer a grill-me interview by voice, then turn the spoken answers into a readback, acceptance criteria, and failing tests you can check without replaying the audio.
  • Three copy-paste prompts that make an agent cope with transcription errors instead of guessing around them.
  • A team checklist, the traps, and the failure modes with fixes.

Which voice input should you use with each agent?

Section titled “Which voice input should you use with each agent?”

Start from the agent you already use. The built-in options differ in kind: Claude Code and Copilot CLI turn speech into prompt text, while Codex CLI holds a spoken conversation.

AgentBuilt-in voiceWhat it doesNeedsChecked against
Claude Code (CLI, agent view, VS Code extension)/voice, /voice hold, /voice tap, /voice offLive transcription into the prompt; you can mix typing and speechclaude.ai login, local microphone, WSLg if you run in WSLClaude Code voice dictation docs, 2026-09-26
Codex CLI/voice, /voice settings, /voice mute, /voice stop, F8Spoken conversation; Codex listens and answers aloudmacOS, an MSVC-based Windows build, or a glibc-based Linux buildcodex-cli 0.157.1 binary and release notes 0.155.0 and 0.156.0
GitHub Copilot CLICtrl+Space, /voice devicesToggles voice dictation; /voice devices picks and remembers the microphoneA voice runtime the CLI downloadsgithub/copilot-cli changelog 1.0.71 and 1.0.81
CursorVoice input in Agent (secondary sources only, see the Cursor tab)Speech into the Agent promptNot verified for this pageNot verified: cursor.com unreachable on 2026-09-26
Aider/voice, then EnterRecords, transcribes, and sends; mature, with slowing releases (last PyPI release 2026-02-12)Aider’s optional voice dependenciesAider usage/voice.md and PyPI aider-chat, 2026-09-26
Any toolmacOS Dictation, Windows voice typing (Win+H)Types into whatever window has focusOS supportOS vendor

Every built-in option in the table streams audio to a vendor. Claude Code’s documentation says: “Audio is not processed locally.” Copilot CLI’s changelog does not say where transcription runs, so treat it as cloud until you confirm otherwise.

Popularity as of 2026-09-26: no vendor publishes usage numbers for voice input, and there is no install count because the feature ships inside each agent (this site’s agent-tools dossier found no public metric). What can be dated is release activity: Claude Code’s changelog lists 20 dictation languages from 2.1.69 (published to npm on 2026-03-04), Copilot CLI added Ctrl+Space dictation in 1.0.81 (2026-08-27), and Codex 0.156.0 (2026-09-22) turned voice conversations on by default.

Set up Claude Code voice dictation in tap mode

Section titled “Set up Claude Code voice dictation in tap mode”

Tap mode suits long answers: you tap once, speak for as long as you need, and tap again to send. Unlike hold mode, it needs no key-repeat warmup, which fails in terminals with key repeat off.

  1. Confirm you are signed in with a claude.ai account. Run /status in Claude Code. If the login method is an API key, Bedrock, Google Cloud’s Agent Platform, or Microsoft Foundry, dictation is not available; run /login and sign in with claude.ai.

  2. Turn on dictation in tap mode:

    /voice tap

    When you enable it, Claude Code runs a microphone check. On macOS the first run triggers the system permission prompt for your terminal app.

  3. With the prompt empty, tap Space. The footer shows ● REC · tap to send. Speak, then tap Space again. Claude Code inserts the transcript and submits it when it is at least three words long, so a stray tap never sends one word. Recording also stops after 15 seconds of silence or two minutes in total. Esc or Ctrl+C discards the recording and restores the prompt.

  4. To dictate in Polish (or another of the 20 supported languages), set language in /config or in ~/.claude/settings.json. The same setting makes Claude answer in that language and name sessions in it:

    {
    "language": "polish",
    "voice": {
    "enabled": true,
    "mode": "tap"
    }
    }

    /voice writes the voice object for you, and it persists across sessions. In hold mode, add "autoSubmit": true to send on key release; without it, you press Enter yourself.

  5. Optional: move dictation off Space. Bind voice:pushToTalk to another key in ~/.claude/keybindings.json; the new key replaces Space (the "space": null line only makes that explicit). A modifier combination such as meta+k starts recording on the first keypress in both modes, with no warmup:

    {
    "bindings": [
    {
    "context": "Chat",
    "bindings": {
    "meta+k": "voice:pushToTalk",
    "space": null
    }
    }
    ]
    }

You should see your words appear dimmed in the prompt as you speak, then turn solid when the transcript is final. Claude Code adds your project name and git branch name as recognition hints and is tuned for terms such as OAuth, JSON, and localhost. Your own table and function names are not on that list, which is why the workflow below checks them.

Transcription does not consume Claude messages or tokens and does not count toward the limits in /usage, according to Claude Code’s documentation. Context cost is the same as for a typed prompt of the same length, and a two-minute spoken answer can add a few hundred words.

Voice input in Codex, Cursor, Copilot CLI, and other tools

Section titled “Voice input in Codex, Cursor, Copilot CLI, and other tools”

Use the setup above. Dictation also works in agent view: hold or tap your push-to-talk key while the dispatch input or a peek-panel reply is focused to dictate to a background session. The VS Code extension supports dictation with the same claude.ai requirement, but not in VS Code Remote sessions (SSH, Dev Containers, and Codespaces), because the microphone is on your machine and the extension runs on the remote host.

Dictate your grill-me answers: the workflow

Section titled “Dictate your grill-me answers: the workflow”

A /grill-me or /grill-with-docs session is the best fit for voice because the agent asks a round of numbered questions, each with a recommended answer, and waits for yours. Your job is to accept, reject, or qualify each recommendation, and the qualification is the part people drop when they have to type it. The risk is that the transcript, not your intent, becomes the spec. The workflow keeps the typing for the parts that must be exact and puts a check between speech and code.

Install the interview skills in your agent

Section titled “Install the interview skills in your agent”

grill-with-docs ships in Matt Pocock’s mattpocock/skills pack. Install it once, with the plugin or the skills CLI, not both: the pack’s README warns that every skill then appears twice. The invocation differs per tool.

Terminal window
# terminal: official marketplace, nothing to add first
claude plugin install mattpocock-skills

Plugin skills are namespaced, so you start the interview with /mattpocock-skills:grill-with-docs. Run /mattpocock-skills:setup-matt-pocock-skills once per repository first; it asks where the glossary and ADRs go. The plugin (1.2.3, 25 skills) adds about 1,600 tokens of always-loaded skill descriptions, measured by this site’s frameworks dossier on 2026-09-26; check yours with /context.

  1. Start the session by typing. Type your tool’s command from the tabs above, a space, and then paste the first prompt below. Commands, file paths, and identifiers are faster and safer typed than spoken.

  2. Answer each question by voice. In Claude Code with /voice tap, tap, speak and tap. The agent asks a numbered round, so start each spoken answer with its number, then use the same four-part shape: the decision, the reason, the constraint, and what would change your mind. “Q3: prorate to the day, because finance reconciles daily, except annual plans, which keep the seat until renewal; I would change this if Stripe cannot do per-day credits.”

  3. Type the identifiers. In tap mode the second tap sends the answer, so it cannot pause for typing. When an answer must name a table, column, function, or flag, switch to /voice hold (without autoSubmit): hold Space and speak, release, type the name, hold again to continue, and press Enter when the answer is complete. Claude Code inserts each transcript at the cursor, so speech and typing mix in one prompt.

  4. Ask for a readback before anything is written. At the end of the interview, use the second prompt below. The agent lists every decision with the words you actually said and flags anything that looks misheard. Read that list, not the transcript.

  5. Turn decisions into things a machine checks. Use the third prompt to write acceptance criteria and failing tests. With /grill-with-docs, the resolved terms land in GLOSSARY.md and hard decisions in ADRs; review those diffs with git diff, because they outlive the conversation. Repositories set up with an older version of the skills may still have CONTEXT.md, the name the skills used before they switched to GLOSSARY.md (mattpocock/skills README, checked 2026-09-26); rename it with git mv.

  6. Sign off on the artifacts, not the audio. You approve the readback, the acceptance criteria, and the red tests. From here the build runs like any spec-driven change: the tests going green is the proof, and nobody replays what you said.

How do you know the agent heard what you meant?

Section titled “How do you know the agent heard what you meant?”

You do not check the transcript word by word. You check three artifacts that a misheard word cannot slip through:

  • The readback list. Every decision in one sentence beside your quoted words. A wrong quote is a transcription error; a right quote with a wrong decision is a misunderstanding. Both surface here, before any code exists.
  • The acceptance criteria and red tests. A misheard “annual” versus “and you’ll” becomes a Given/When/Then line you can read in seconds, and a test that fails for the stated reason proves the criterion is executable.
  • The GLOSSARY.md and ADR diffs from /grill-with-docs. Terms in the glossary must match the names in the code. A term that exists nowhere in the repository is almost always a transcription artifact.

The developer who ran the interview signs off on the readback and the red tests. On a team, the reviewer of the pull request reads the acceptance criteria and the ADRs, not the conversation log.

For a tech lead, the questions are data handling and where it works, not whether people like it.

CheckWhy it mattersWhat to do
Where does audio go?Claude Code streams audio to Anthropic for transcription; Codex voice is a live conversation with OpenAI’s serviceTreat speech like prompt text under your AI data policy; read Claude Code’s data-usage page
Which sign-in does your team use?Claude Code dictation needs a claude.ai login; API-key, Bedrock, Agent Platform, and Foundry setups cannot use itIf you standardised on a cloud provider, plan on system dictation instead
Where do people work?No dictation over SSH, in cloud sessions, or in VS Code Remote, Dev Containers, or CodespacesRemote-first setups use system dictation on the local machine
Can it be turned off?An administrator policy can disable Claude Code voice dictation; users then see Voice mode is disabled by your organization's policyDecide before rollout, not after the first complaint
Shared officesSpoken prompts include customer names and incident detailsHeadsets and push-to-talk (hold mode) in open-plan spaces

What breaks with voice dictation, and how to fix it

Section titled “What breaks with voice dictation, and how to fix it”
SymptomCauseFix
Voice mode requires a Claude.ai accountSigned in with an API key or a third-party provider/login with claude.ai, or use system dictation
Holding Space types spaces and nothing recordsDictation is off, or the terminal sends no key-repeat events/voice hold to turn it on; if one or two spaces appear and stop, switch to /voice tap
Tapping Space types a space in tap modeThe first tap only records when the prompt is emptyClear the prompt, or rebind to meta+k
Microphone access is deniedThe terminal app lacks microphone permissionmacOS: System Settings › Privacy & Security › Microphone; Windows: allow microphone access for desktop apps; run /voice again
Terminal missing from the macOS microphone listStale permission statetccutil reset Microphone com.apple.Terminal (or your terminal’s bundle ID), then quit (Cmd+Q) and relaunch the terminal
Voice mode requires SoX for audio recording on LinuxNative audio module did not load and no fallback existsInstall SoX with the command the error shows
could not find a working audio recorder in WSLSoX lacks its PulseAudio backendsudo apt install sox libsox-fmt-pulse
Voice input is failing repeatedly and has been pausedThree failures in 10 seconds, usually no capture device (headless host, remote shell)Fix the device, wait 10 seconds, try again
Voice stream error: WebSocket upgrade rejected with HTTP 4xxStale sign-in, or a proxy or bot-protection service answering instead/login; check the VPN or proxy on the path
Garbled text or the wrong languageDictation defaults to EnglishSet language in /config first
Codex /voice refuses to startNo conversation yet, a side conversation, or an unsupported buildSend a first message or return to the main thread. Unsupported build: codex-cli 0.157.1 on Linux x64 from npm is musl (see the Codex tab)

If a misheard answer already reached the agent, do not argue with it in the next turn. In Claude Code, press Esc twice with an empty prompt to open the rewind menu and restore the conversation before that answer, or type the correction explicitly (“Correction to decision 4: the column is seat_count, not seed count”) and ask for the readback again.

Where to go next with voice-driven planning

Section titled “Where to go next with voice-driven planning”

Frequently asked questions

How do I dictate prompts in Claude Code?

Run /voice to turn dictation on, then hold Space while you speak, or run /voice tap to tap Space once to start and again to send. You need a claude.ai login and a local microphone; it does not work with an API key, Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, SSH, or cloud sessions.

Can I dictate to Claude Code in Polish?

Yes. Dictation supports 20 languages including Polish. It follows the language setting, which also sets the language Claude answers in, so setting it to Polish changes both.

Does Codex CLI have voice dictation?

Not as prompt dictation. In codex-cli 0.157.1, /voice in the TUI starts a spoken voice conversation rather than dictation into the prompt; the in_app_dictation feature flag is on, but codex-cli 0.157.1 exposes no TUI command for it (checked 2026-09-26). /voice starts or stops the conversation, /voice settings picks a voice, and F8 toggles it. It requires macOS, an MSVC-based Windows build, or a glibc-based Linux build.

Does voice dictation cost tokens?

Not in Claude Code: its documentation says transcription does not consume Claude messages or tokens and does not count toward the /usage limits. The prompt the transcript becomes is billed like any typed prompt.