The gist
Typing your prompts (your requests to the AI) by hand isn't the only option. There are three ways to talk to Claude Code with your voice: standalone dictation apps (for example, Aqua Voice), the built-in /voice mode, and local Whisper for people who don't want to send audio to the cloud. When your hands are busy or your thoughts come faster than your fingers, voice wins.
Key concepts
- Aqua Voice: a specialized STT tool (Speech-to-Text, speech recognition) focused on technical terms
- Native
/voicemode: Claude Code's built-in voice dictation - Whisper API (application programming interface) from OpenAI: cloud transcription for general use
- Local Whisper: transcription without sending anything to the cloud
- VoiceMode MCP (Model Context Protocol): Whisper STT + Kokoro TTS (Text-to-Speech, turning text into speech) for two-way voice conversations
- Dictation in any app: Aqua Voice, Wispr Flow, MacWhisper: you speak, and the text appears wherever your cursor is
Theory
Aqua Voice: why it's built for developers
Ordinary voice assistants (Siri, Google dictation) struggle with technical terms:
claude-sonnet-5-5→ "clawed sonnet five five"npm install -g→ "NPM in stall G"Anthropic API key→ "and the topic API key" (good luck guessing)
Aqua Voice is tuned for technical terms, bash commands (bash is the terminal's command language), library names and code patterns, according to its developers.
Specs (as of October 2026, from the service's website):
| Parameter | Value |
|---|---|
| Purpose | Dictation focused on technical terms |
| Accuracy | The service claims 97.3% on its own AISpeak benchmark (vendor data, not independently verified) |
| Languages | 49 claimed; if you dictate in a language other than English, check support on the website before buying |
| Integration | System-wide (works everywhere) |
| Platforms | The main platform is macOS; check the website for other versions |
| Price | There's a free tier with a word limit and paid plans with unlimited words; current prices are on the website |
Alternatives (as of October 2026): Wispr Flow (Mac, Windows, iPhone and Android, 100+ languages, a free tier with a weekly word limit) and MacWhisper (runs locally on Mac, 100+ languages, the core features are free). Catalog and current terms: Tools.
How it works:
- Press the hotkey (you can configure it)
- Speak, and you see a waveform in the menu bar
- Release the key → the text is inserted wherever your cursor is
- Works in VS Code, Terminal, the browser, any text field
Installation:
# From the website
# aquavoice.com → Get Started on Mac → download and installFirst-time setup:
- Launch Aqua Voice
- System Settings → Privacy & Security → Microphone → allow
- Set a hotkey (I recommend Option+Space or Cmd+Shift+V)
- Pick the mode that fits your work (see the app settings for the list of modes)
Native /voice mode in Claude Code
Claude Code has built-in voice dictation, with no third-party apps.
Turning it on:
# Type this in Claude Code /voice # Choose a mode: hold the key, or tap to start and tap to stop /voice hold /voice tap
What it does:
- Hold mode (the default): hold Space → speak → release → the text is inserted into your request
- Tap mode: press Space → speak → press it again → the request is sent
- Transcription is tuned for programming terms (regex, OAuth, JSON), and your project name and git branch are suggested automatically
- Works in the terminal CLI and in the VS Code extension; doesn't work in cloud sessions or over SSH
- Requires signing in with a claude.ai account (not an API key); transcription doesn't use tokens or your plan's usage limits
- The dictation language comes from the
languagesetting (/config); the default is English, and other languages are supported: if you dictate in another language, select it, or the text will come out garbled - Audio is sent to Anthropic's servers for transcription: for confidential recordings, use local Whisper
How it differs from Aqua Voice:
| Parameter | /voice mode |
Aqua Voice |
|---|---|---|
| Availability | Only in Claude Code (CLI and VS Code) | System-wide (everywhere) |
| Setup | /voice and picking a language |
5 minutes |
| Technical accuracy | Tuned for programming terms | 97.3% claimed (vendor data) |
| Price | Doesn't use tokens or plan limits | Free tier and paid plans |
| Customization | Hold/tap modes, configurable key | Hotkeys, modes |
| Offline | No (audio goes to Anthropic's servers) | Check the website |
Recommendation: /voice for a quick start and occasional use. Aqua Voice if voice becomes your main way of typing.
Whisper API: cloud transcription
OpenAI Whisper is a popular transcription model. Through the API it's available as the whisper-1 model, billed per minute of audio. OpenAI also has newer speech recognition models (for example, gpt-4o-transcribe and gpt-4o-mini-transcribe). For current prices, see OpenAI's pricing page.
Using it from Python (a programming language):
import openai
import sounddevice as sd
import numpy as np
import scipy.io.wavfile as wav
import tempfile
client = openai.OpenAI() # the key is read from the OPENAI_API_KEY environment variable
def record_and_transcribe(duration_seconds=10):
"""Records audio and transcribes it with Whisper"""
print(f"Recording {duration_seconds} seconds... (start talking!)")
# Recording
sample_rate = 16000
recording = sd.rec(
int(duration_seconds * sample_rate),
samplerate=sample_rate,
channels=1
)
sd.wait()
print("Transcribing...")
# Save to a temporary file
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:
wav.write(tmp.name, sample_rate, recording)
with open(tmp.name, "rb") as audio_file:
transcription = client.audio.transcriptions.create(
model="whisper-1",
file=audio_file,
language="en" # or "es", or remove it for auto-detection
)
return transcription.text
# Example of using it with Claude
import anthropic
claude = anthropic.Anthropic() # the key is read from the ANTHROPIC_API_KEY environment variable
while True:
user_input = record_and_transcribe(duration_seconds=8)
print(f"You: {user_input}")
if "stop" in user_input.lower():
break
response = claude.messages.create(
model="claude-sonnet-5-5",
max_tokens=1000,
messages=[{"role": "user", "content": user_input}]
)
print(f"Claude: {''.join(b.text for b in response.content if b.type == 'text')}")Installing the dependencies:
pip install openai sounddevice scipy numpy anthropicLocal Whisper: full privacy
For work where you can't send audio to the cloud (client negotiations, confidential data).
# Installation
pip install openai-whisper ffmpeg-python
# On a Mac: brew install ffmpegimport whisper
# Load the model (once; it gets cached)
# tiny=39MB, base=74MB, small=244MB, medium=769MB, large=1.5GB
model = whisper.load_model("small") # balance of speed and quality
def transcribe_local(audio_path: str, language: str = "en") -> str:
"""Transcribes audio locally, without the cloud"""
result = model.transcribe(
audio_path,
language=language,
fp16=False # fp16=True if you have a GPU
)
return result["text"]
# Example
text = transcribe_local("recording.wav")
print(text)Comparing Whisper models:
| Model | Size | Speed (CPU, approximate) | Quality |
|---|---|---|---|
| tiny | 39 MB | ~10x real time | Basic |
| base | 74 MB | ~7x real time | Decent |
| small | 244 MB | ~4x real time | Good |
| medium | 769 MB | ~2x real time | Very good |
| large | 1.5 GB | ~1x real time | Best |
On Apple Silicon, optimized implementations run faster (for example, whisper.cpp or the MacWhisper app); the standard openai-whisper package runs on CPU or GPU.
VoiceMode MCP: two-way voice
VoiceMode MCP lets Claude Code talk back, not just listen.
# Install as a Claude Code plugin
claude plugin install voicemode@voicemode
/voicemode:install
/voicemode:converse
# Alternative: the Python installer
uvx voice-mode-installWhat it does:
- STT: Whisper.cpp (local), with the OpenAI API as a fallback
- TTS: Kokoro TTS (open source, runs locally) or the OpenAI API
- Integration with Claude Code for voice conversations
Useful for: pair programming with AI, code reviews out loud, dictating documentation.
Practical scenarios
Scenario 1: Coding by voice
Open VS Code, turn on Aqua Voice
→ Say: "add email validation to the signup function"
→ The text goes into the Claude Code chat
→ Claude writes the code
→ Say the next task out loudScenario 2: A developer's voice journal
#!/usr/bin/env bash
# voice-standup.sh — a voice standup every morning
RECORDING_FILE="/tmp/standup-$(date +%Y%m%d).wav"
# Record your voice (90 seconds)
rec -r 16000 -c 1 "$RECORDING_FILE" trim 0 90
# Transcribe with the Whisper API
TRANSCRIPT=$(curl -s -X POST "https://api.openai.com/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file="@$RECORDING_FILE" \
-F model="whisper-1" \
-F language="en" | jq -r '.text')
# Structure it with Claude
STANDUP=$(echo "$TRANSCRIPT" | claude -p "
Turn this voice standup into a structured format:
- What I did yesterday
- What I plan to do today
- Blockers
")
# Save it
echo "$STANDUP" >> ~/standups/$(date +%Y-%m).md
echo "Standup saved!"Scenario 3: Transcribing client meetings
# Local Whisper: the audio doesn't go to the cloud (the transcript text then goes to the Claude API)
# Only record conversations with everyone's consent
import whisper
import anthropic
model = whisper.load_model("medium")
claude = anthropic.Anthropic()
# Transcribe the meeting recording
transcript = model.transcribe("meeting-2026-05-09.mp3", language="en")
# Extract tasks with Claude
response = claude.messages.create(
model="claude-sonnet-5-5",
max_tokens=2000,
messages=[{
"role": "user",
"content": f"From this meeting transcript, extract: "
f"a list of tasks with owners, key decisions, "
f"and open questions.\n\nTranscript:\n{transcript['text']}"
}]
)
print("".join(b.text for b in response.content if b.type == "text"))macOS Dictation: the free alternative
The dictation feature built into macOS (in many cases it works without internet, depending on the language and the macOS version):
System Settings → Keyboard → Dictation → On
Hotkey: press Fn twice (or set your own)
Language: English (many other languages are supported)Pros: free, often offline, works everywhere Cons: worse with technical terms, no API
Practice
- Turn on macOS Dictation (System Settings → Keyboard) and try dictating a prompt into Claude Code
- Install and set up Aqua Voice (aquavoice.com) or another dictation option, and compare the accuracy on technical terms
- Create a
voice-note.shscript that records a voice note and saves it to a Markdown file - (optional) Connect VoiceMode MCP for two-way voice conversations
Tools and resources
- Aqua Voice: specialized STT for developers
- Claude Code: voice dictation: documentation for
/voice - Wispr Flow and MacWhisper: dictation alternatives, see the Tools catalog
- OpenAI Whisper API: cloud transcription
- OpenAI Whisper (local): GitHub repository
- VoiceMode MCP: two-way voice for Claude
- Kokoro TTS: open-source text-to-speech
- macOS Dictation: System Settings → Keyboard → Dictation
Key takeaways
Aqua Voice is one option for technical voice input: it works system-wide in every app. Check the website for prices and language support.
Native
/voicemode in Claude Code takes almost no setup, doesn't use tokens, and requires a claude.ai sign-in. For occasional use, it's enough. For constant input across all your apps, get a standalone dictation app.
Local Whisper is for private data. Transcripts of client negotiations and internal meetings all stay on your machine.
Next lesson
→ Video editing with Claude Code: FFmpeg, DaVinci Resolve MCP, Remotion
The mark stays in this browser only and is never sent anywhere. My progress