Library · Power-user techniques

Voice tools: Aqua Voice, /voice mode, Whisper

Confident user40 minUpdated: October 2026
48 of 105 in the library

Time: ~20 min theory + 20 min practice


The gist

Typing your prompts (your requests to the AI) by hand isn't the only option. There are three ways to talk to Claude Code with your voice: standalone dictation apps (for example, Aqua Voice), the built-in /voice mode, and local Whisper for people who don't want to send audio to the cloud. When your hands are busy or your thoughts come faster than your fingers, voice wins.

🎨 Picture this: voice input in Claude Code is like dictating a letter to an assistant. You talk, they write it down. The difference is that an assistant draws a salary, while a dictation app costs a subscription or nothing at all. And it never gets tired.


Key concepts

  • Aqua Voice: a specialized STT tool (Speech-to-Text, speech recognition) focused on technical terms
  • Native /voice mode: Claude Code's built-in voice dictation
  • Whisper API (application programming interface) from OpenAI: cloud transcription for general use
  • Local Whisper: transcription without sending anything to the cloud
  • VoiceMode MCP (Model Context Protocol): Whisper STT + Kokoro TTS (Text-to-Speech, turning text into speech) for two-way voice conversations
  • Dictation in any app: Aqua Voice, Wispr Flow, MacWhisper: you speak, and the text appears wherever your cursor is

Theory

Aqua Voice: why it's built for developers

Ordinary voice assistants (Siri, Google dictation) struggle with technical terms:

  • claude-sonnet-5-5 → "clawed sonnet five five"
  • npm install -g → "NPM in stall G"
  • Anthropic API key → "and the topic API key" (good luck guessing)

🎨 Picture this: an ordinary dictation tool is like an interpreter who only knows everyday words. Fine for grocery shopping, a disaster in a courtroom. Aqua Voice is like an interpreter who spent three years in Silicon Valley and knows all the technical jargon.

Aqua Voice is tuned for technical terms, bash commands (bash is the terminal's command language), library names and code patterns, according to its developers.

Specs (as of October 2026, from the service's website):

Parameter Value
Purpose Dictation focused on technical terms
Accuracy The service claims 97.3% on its own AISpeak benchmark (vendor data, not independently verified)
Languages 49 claimed; if you dictate in a language other than English, check support on the website before buying
Integration System-wide (works everywhere)
Platforms The main platform is macOS; check the website for other versions
Price There's a free tier with a word limit and paid plans with unlimited words; current prices are on the website

Alternatives (as of October 2026): Wispr Flow (Mac, Windows, iPhone and Android, 100+ languages, a free tier with a weekly word limit) and MacWhisper (runs locally on Mac, 100+ languages, the core features are free). Catalog and current terms: Tools.

How it works:

  1. Press the hotkey (you can configure it)
  2. Speak, and you see a waveform in the menu bar
  3. Release the key → the text is inserted wherever your cursor is
  4. Works in VS Code, Terminal, the browser, any text field

Installation:

bash
# From the website
# aquavoice.com → Get Started on Mac → download and install

First-time setup:

  1. Launch Aqua Voice
  2. System Settings → Privacy & Security → Microphone → allow
  3. Set a hotkey (I recommend Option+Space or Cmd+Shift+V)
  4. Pick the mode that fits your work (see the app settings for the list of modes)

Native /voice mode in Claude Code

Claude Code has built-in voice dictation, with no third-party apps.

🎨 Picture this: you used to have to plug in an external microphone through an adapter. Now the microphone is built in: just open it and talk.

Turning it on:

Type this into the chat
# Type this in Claude Code
/voice

# Choose a mode: hold the key, or tap to start and tap to stop
/voice hold
/voice tap

What it does:

  • Hold mode (the default): hold Space → speak → release → the text is inserted into your request
  • Tap mode: press Space → speak → press it again → the request is sent
  • Transcription is tuned for programming terms (regex, OAuth, JSON), and your project name and git branch are suggested automatically
  • Works in the terminal CLI and in the VS Code extension; doesn't work in cloud sessions or over SSH
  • Requires signing in with a claude.ai account (not an API key); transcription doesn't use tokens or your plan's usage limits
  • The dictation language comes from the language setting (/config); the default is English, and other languages are supported: if you dictate in another language, select it, or the text will come out garbled
  • Audio is sent to Anthropic's servers for transcription: for confidential recordings, use local Whisper

How it differs from Aqua Voice:

Parameter /voice mode Aqua Voice
Availability Only in Claude Code (CLI and VS Code) System-wide (everywhere)
Setup /voice and picking a language 5 minutes
Technical accuracy Tuned for programming terms 97.3% claimed (vendor data)
Price Doesn't use tokens or plan limits Free tier and paid plans
Customization Hold/tap modes, configurable key Hotkeys, modes
Offline No (audio goes to Anthropic's servers) Check the website

Recommendation: /voice for a quick start and occasional use. Aqua Voice if voice becomes your main way of typing.


Whisper API: cloud transcription

OpenAI Whisper is a popular transcription model. Through the API it's available as the whisper-1 model, billed per minute of audio. OpenAI also has newer speech recognition models (for example, gpt-4o-transcribe and gpt-4o-mini-transcribe). For current prices, see OpenAI's pricing page.

🎨 Picture this: Whisper is like an experienced court interpreter. Accurate, multilingual, handles recordings of varying quality. But every time, you have to send them the recording in the cloud.

Using it from Python (a programming language):

python
import openai
import sounddevice as sd
import numpy as np
import scipy.io.wavfile as wav
import tempfile

client = openai.OpenAI()  # the key is read from the OPENAI_API_KEY environment variable

def record_and_transcribe(duration_seconds=10):
    """Records audio and transcribes it with Whisper"""
    print(f"Recording {duration_seconds} seconds... (start talking!)")
    
    # Recording
    sample_rate = 16000
    recording = sd.rec(
        int(duration_seconds * sample_rate),
        samplerate=sample_rate,
        channels=1
    )
    sd.wait()
    
    print("Transcribing...")
    
    # Save to a temporary file
    with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:
        wav.write(tmp.name, sample_rate, recording)
        
        with open(tmp.name, "rb") as audio_file:
            transcription = client.audio.transcriptions.create(
                model="whisper-1",
                file=audio_file,
                language="en"  # or "es", or remove it for auto-detection
            )
    
    return transcription.text

# Example of using it with Claude
import anthropic

claude = anthropic.Anthropic()  # the key is read from the ANTHROPIC_API_KEY environment variable

while True:
    user_input = record_and_transcribe(duration_seconds=8)
    print(f"You: {user_input}")
    
    if "stop" in user_input.lower():
        break
    
    response = claude.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1000,
        messages=[{"role": "user", "content": user_input}]
    )
    
    print(f"Claude: {''.join(b.text for b in response.content if b.type == 'text')}")

Installing the dependencies:

bash
pip install openai sounddevice scipy numpy anthropic

Local Whisper: full privacy

For work where you can't send audio to the cloud (client negotiations, confidential data).

🎨 Picture this: local Whisper is like hiring an interpreter who comes to your home and signs an NDA. Slower than the cloud, but everything stays with you.

bash
# Installation
pip install openai-whisper ffmpeg-python
# On a Mac: brew install ffmpeg
python
import whisper

# Load the model (once; it gets cached)
# tiny=39MB, base=74MB, small=244MB, medium=769MB, large=1.5GB
model = whisper.load_model("small")  # balance of speed and quality

def transcribe_local(audio_path: str, language: str = "en") -> str:
    """Transcribes audio locally, without the cloud"""
    result = model.transcribe(
        audio_path,
        language=language,
        fp16=False  # fp16=True if you have a GPU
    )
    return result["text"]

# Example
text = transcribe_local("recording.wav")
print(text)

Comparing Whisper models:

Model Size Speed (CPU, approximate) Quality
tiny 39 MB ~10x real time Basic
base 74 MB ~7x real time Decent
small 244 MB ~4x real time Good
medium 769 MB ~2x real time Very good
large 1.5 GB ~1x real time Best

On Apple Silicon, optimized implementations run faster (for example, whisper.cpp or the MacWhisper app); the standard openai-whisper package runs on CPU or GPU.


VoiceMode MCP: two-way voice

VoiceMode MCP lets Claude Code talk back, not just listen.

🎨 Picture this: without VoiceMode, it's a conversation through a glass wall. You talk, it replies in text. With VoiceMode, it's a normal out-loud conversation, like with a coworker.

bash
# Install as a Claude Code plugin
claude plugin install voicemode@voicemode
/voicemode:install
/voicemode:converse

# Alternative: the Python installer
uvx voice-mode-install

What it does:

  • STT: Whisper.cpp (local), with the OpenAI API as a fallback
  • TTS: Kokoro TTS (open source, runs locally) or the OpenAI API
  • Integration with Claude Code for voice conversations

Useful for: pair programming with AI, code reviews out loud, dictating documentation.


Practical scenarios

Scenario 1: Coding by voice

Code
Open VS Code, turn on Aqua Voice
→ Say: "add email validation to the signup function"
→ The text goes into the Claude Code chat
→ Claude writes the code
→ Say the next task out loud

Scenario 2: A developer's voice journal

bash
#!/usr/bin/env bash
# voice-standup.sh — a voice standup every morning

RECORDING_FILE="/tmp/standup-$(date +%Y%m%d).wav"

# Record your voice (90 seconds)
rec -r 16000 -c 1 "$RECORDING_FILE" trim 0 90

# Transcribe with the Whisper API
TRANSCRIPT=$(curl -s -X POST "https://api.openai.com/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file="@$RECORDING_FILE" \
  -F model="whisper-1" \
  -F language="en" | jq -r '.text')

# Structure it with Claude
STANDUP=$(echo "$TRANSCRIPT" | claude -p "
Turn this voice standup into a structured format:
- What I did yesterday
- What I plan to do today
- Blockers
")

# Save it
echo "$STANDUP" >> ~/standups/$(date +%Y-%m).md
echo "Standup saved!"

Scenario 3: Transcribing client meetings

python
# Local Whisper: the audio doesn't go to the cloud (the transcript text then goes to the Claude API)
# Only record conversations with everyone's consent
import whisper
import anthropic

model = whisper.load_model("medium")
claude = anthropic.Anthropic()

# Transcribe the meeting recording
transcript = model.transcribe("meeting-2026-05-09.mp3", language="en")

# Extract tasks with Claude
response = claude.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2000,
    messages=[{
        "role": "user",
        "content": f"From this meeting transcript, extract: "
                   f"a list of tasks with owners, key decisions, "
                   f"and open questions.\n\nTranscript:\n{transcript['text']}"
    }]
)

print("".join(b.text for b in response.content if b.type == "text"))

macOS Dictation: the free alternative

The dictation feature built into macOS (in many cases it works without internet, depending on the language and the macOS version):

Code
System Settings → Keyboard → Dictation → On
Hotkey: press Fn twice (or set your own)
Language: English (many other languages are supported)

Pros: free, often offline, works everywhere Cons: worse with technical terms, no API


Practice

  1. Turn on macOS Dictation (System Settings → Keyboard) and try dictating a prompt into Claude Code
  2. Install and set up Aqua Voice (aquavoice.com) or another dictation option, and compare the accuracy on technical terms
  3. Create a voice-note.sh script that records a voice note and saves it to a Markdown file
  4. (optional) Connect VoiceMode MCP for two-way voice conversations

Tools and resources


Key takeaways

Aqua Voice is one option for technical voice input: it works system-wide in every app. Check the website for prices and language support.

Native /voice mode in Claude Code takes almost no setup, doesn't use tokens, and requires a claude.ai sign-in. For occasional use, it's enough. For constant input across all your apps, get a standalone dictation app.

Local Whisper is for private data. Transcripts of client negotiations and internal meetings all stay on your machine.


Next lesson

→ Video editing with Claude Code: FFmpeg, DaVinci Resolve MCP, Remotion

The mark stays in this browser only and is never sent anywhere. My progress