Library · Content studio: design, voice, video, social

AI voiceover: ElevenLabs, the TTS stack, voice content

Builder60 minUpdated: October 2026
52 of 105 in the library

Module: Content Factory AI | Time: about 25 min theory + 35 min practice


The gist

A voice actor is a separate line item: scheduling, a studio, an hourly rate. A speech synthesis platform is billed differently: you pay for volume (credits per month), and on long texts the difference is noticeable (current plans: What's current). Claude writes the script (a narration script or program code), ElevenLabs voices it, FFmpeg stitches the video together. No voice actor, no studio, no scheduling. Set up the pipeline (a chain of sequential steps) once, and from then on the factory runs by itself.

🎨 Picture this: a voice actor is like a cab: you pay for every ride. A TTS (Text-to-Speech, synthesizing speech from text) pipeline is like owning a car: you buy it once and drive as much as you want. And every additional ride costs pennies.


Key concepts

  • TTS (Text-to-Speech): the technology for synthesizing speech from text
  • Voice cloning: creating an AI copy of a voice, a digital clone of your voice from a sample
  • ElevenLabs: one of the leading commercial TTS platforms
  • ElevenLabs MCP (Model Context Protocol): a direct integration with Claude Code: a hosted server with OAuth sign-in, or a local server
  • Eleven v4: ElevenLabs' new model (September 28, 2026): more control over intonation, 90+ languages; before it, Multilingual v2 was the main one
  • Kokoro TTS: an open source alternative for offline voiceover (a limited set of languages, see below)
  • TTS pipeline: Claude writes the script → TTS voices it → FFmpeg assembles the video

Theory

Why TTS belongs in an AI stack: a question of scale

Voiceover is the last manual step in content production. Claude writes text automatically, Midjourney generates images, FFmpeg assembles video. But the narrator is still a live person with a schedule and a price tag.

TTS closes that gap.

Three scenarios where TTS changes the economics:

Scenario 1: A YouTube channel at scale

One narrator can voice 4-5 videos a week: that's the physical limit. With TTS it's 20-30 videos of the same quality in a single overnight run. The voice is always in the same mood, never gets tired and never asks for a retake.

Scenario 2: Localization

Translating a video into 10 languages with live narrators means 10 narrators and 10 schedules. With Eleven v4, the same voice speaks dozens of languages (90+ according to ElevenLabs as of October 2026), and the pipeline itself takes hours. A native speaker checks pronunciation and translation before publishing.

Scenario 3: Audiobooks and podcasts

An 80,000-word book takes a narrator 8-10 hours of recording plus editing. Synthesis uses up your plan's credits: count the characters in the manuscript and compare that with your plan's limit (current prices: What's current). Claude reads the manuscript, adapts it for listening, and ElevenLabs voices it chapter by chapter.


ElevenLabs: how it works and how it's priced

ElevenLabs is one of the best-known professional TTS platforms.

What's inside (as of October 2026):

Parameter Value
Models Eleven v4 (released September 28, 2026), previously Multilingual v2 was the main one; there are fast Flash and Turbo models
Languages 90+ for Eleven v4 (according to ElevenLabs)
Output formats several, set with the output_format parameter (for example, mp3_44100_128)
Cloning instant clone and professional clone
Beyond voiceover speech recognition, sound effects, music, voice agents

Plans as of October 2026 (current prices: What's current and the ElevenLabs pricing page):

Plan Price per month Credits per month Commercial license Cloning
Free $0 10,000 no no
Starter $6 see the pricing page yes instant clone
Creator $22 (first month $11) see the pricing page yes instant and professional clone
Pro $99 see the pricing page yes instant and professional clone

Credit usage depends on the model and the length of the text: measure it on a short test and multiply by your volume. The free plan has no commercial license, so for client projects you need a paid one.


Voice cloning: a photograph of your voice

Voice cloning is creating a digital clone of a voice from an audio sample. Once it's cloned, the system synthesizes speech with the same timbre, intonation and rhythm.

🎨 Picture this: voice cloning is like a photograph. You sit in front of the camera once, and after that your image can be printed thousands of times. The voice is recorded once and can be used for voiceover endlessly.

Sample requirements:

  • Instant clone: a short sample of clean speech (ElevenLabs says Eleven v4 can clone from 10 seconds of recording). The cleaner and more varied the recording, the better the result
  • Professional clone: a much longer recording (see the ElevenLabs help center for exact requirements)
  • Format: WAV or MP3, no background noise
  • No music, no other voices, no echo

Instant cloning, step by step:

  1. Go to elevenlabs.io → the Voices section → add a voice → instant clone (requires Starter or higher)
  2. Upload the audio file
  3. Give the voice a name (for example "MyVoice_EN")
  4. Confirm that the voice is yours or that you have its owner's permission
  5. Wait for processing and copy the Voice ID from the voice settings

After cloning, the voice is available through the API. You need the Voice ID for every request.

Important: ElevenLabs requires the voice owner's consent for cloning. Don't clone voices without permission: it violates the ToS (Terms of Service) and the laws of many countries.


A basic example: Claude writes the script, ElevenLabs voices it

python
from elevenlabs.client import ElevenLabs
import anthropic
import os

anthropic_client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

# Step 1: generate the script with Claude
response = anthropic_client.messages.create(
    model="claude-sonnet-5-5",  # current model IDs: see the Anthropic documentation
    max_tokens=1000,
    messages=[{
        "role": "user",
        "content": (
            "Write a 60-second script for an ad for AI consulting services. "
            "Tone: professional, confident, no empty promises. "
            "About 150-160 words. Only the text to be read aloud, no stage directions."
        )
    }]
)

script = "".join(b.text for b in response.content if b.type == "text")
print(f"Script ({len(script.split())} words):\n{script}\n")

# Step 2: voice it with ElevenLabs
audio = eleven.text_to_speech.convert(
    text=script,
    voice_id="JBFqnCBsd6RMkjVDRZzb",  # a voice from the docs example; get your own Voice ID in the Voices section
    model_id="eleven_v4",             # model as of October 2026; check the documentation for current model names
    output_format="mp3_44100_128",
)

with open("ad-script.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)
print("File saved: ad-script.mp3")

Installation:

bash
pip install elevenlabs anthropic

Environment variables:

bash
export ANTHROPIC_API_KEY="sk-ant-..."
export ELEVENLABS_API_KEY="sk_..."

ElevenLabs MCP: voiceover straight from Claude Code

ElevenLabs offers official MCP servers. As of October 2026, the recommended option is the hosted server: nothing to install on your computer, and you sign in with OAuth. The older local server in the ElevenLabs repository is marked as deprecated. Once connected, Claude can voice text, pick voices and save files right from the chat, with no Python scripts (program code).

Connecting:

bash
# ElevenLabs hosted server (recommended; OAuth sign-in, no API key needed)
claude mcp add --transport http elevenlabs https://api.elevenlabs.io/v1/mcp
# then inside Claude Code: /mcp and complete the sign-in

# The older local server (marked deprecated in the repository; requires uv and an API key in the environment)
claude mcp add elevenlabs --env ELEVENLABS_API_KEY=your_key -- uvx elevenlabs-mcp

Or through .mcp.json (local server):

json
{
  "mcpServers": {
    "elevenlabs": {
      "command": "uvx",
      "args": ["elevenlabs-mcp"],
      "env": {
        "ELEVENLABS_API_KEY": "${ELEVENLABS_API_KEY}"
      }
    }
  }
}

Once connected, Claude gets tools (the set depends on the server; see elevenlabs.io/mcp and the server's repository for the exact list):

  • Text-to-speech synthesis
  • Voice cloning and voice management
  • Transcription, audio cleanup, speech-to-speech conversion
  • Generating soundscapes and music

An example request in the chat:

Type this into the chat
Voice the following text with one of the available voices and save it to intro.mp3:

"Welcome to the lesson on voiceover.
Over the next ten minutes you'll build your first pipeline:
a script, a voice and a finished audio file."

Claude will call the MCP tool and save the file with no extra code.


Free alternatives: when ElevenLabs is overkill

Service Quality Price Cloning Offline
ElevenLabs ⭐⭐⭐⭐⭐ Free and paid plans (see above) ✅ on paid plans ❌
OpenAI TTS ⭐⭐⭐⭐ pay by volume through the API (prices: What's current) ❌ ❌
Kokoro TTS ⭐⭐⭐ Free ❌ (preset voices) ✅
macOS say ⭐⭐ Free ❌ ✅

OpenAI TTS: the API works like other OpenAI endpoints (API access points), comes with 13 built-in voices, and handles English and many other languages (according to OpenAI's documentation). No cloning of your own voice. A good fit when you already use OpenAI.

python
import openai, os

client = openai.OpenAI(api_key=os.environ["OPENAI_API_KEY"])

with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",    # tts-1 and tts-1-hd are earlier models
    voice="coral",              # one of the built-in voices, see the list in OpenAI's documentation
    input="Your voiceover text goes here",
    instructions="Speak calmly and in a friendly way.",
) as response:
    response.stream_to_file("output.mp3")

Kokoro TTS: open source, runs locally, weights under the Apache 2.0 license. The quality is lower than ElevenLabs, but it's free and sends no data to the cloud. Note: according to the project README, Kokoro supports American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Chinese. For other languages, Kokoro won't work.

bash
pip install "kokoro>=0.9.4" soundfile
python
from kokoro import KPipeline
import soundfile as sf

pipeline = KPipeline(lang_code="a")  # "a" is American English
generator = pipeline("Text for local voiceover", voice="af_heart")
for i, (_, _, audio) in enumerate(generator):
    sf.write(f"output-{i}.wav", audio, 24000)

macOS say: built into every Mac, free, works offline. The quality is robotic, but for prototypes and testing it's perfect.

bash
# Basic voiceover
say "Hi, this is a test voiceover"

# Save to a file
say -v "Samantha" -o output.aiff "Text to read aloud"

# Convert to MP3 with ffmpeg
say -v "Samantha" -o /tmp/output.aiff "Text" && \
ffmpeg -i /tmp/output.aiff output.mp3 -y -loglevel error

# List the available US English voices
say -v "?" | grep en_US

A YouTube pipeline with TTS voiceover

The full script: a video topic goes in, a finished MP4 (video file) comes out.

bash
#!/usr/bin/env bash
# tts-video-pipeline.sh
# Usage: ./tts-video-pipeline.sh "Video topic" background.jpg
set -euo pipefail

TOPIC="$1"
BACKGROUND="${2:-background.jpg}"
OUTPUT_DIR="./output"
mkdir -p "$OUTPUT_DIR"

echo "Topic: $TOPIC"

# Step 1: Claude writes the script (via the claude CLI)
echo "Writing the script..."
SCRIPT=$(claude -p "Write a 3-minute YouTube script on: $TOPIC.
Tone: conversational, specific, no clichés.
Length: 400-450 words. Only the narrator's text, no stage directions or headings.")

echo "$SCRIPT" > "$OUTPUT_DIR/script.txt"
echo "Script: $(echo "$SCRIPT" | wc -w) words"

# Step 2: ElevenLabs voices it
echo "Recording the voiceover..."
python3 << PYEOF
from elevenlabs.client import ElevenLabs
import os

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

with open("$OUTPUT_DIR/script.txt", encoding="utf-8") as f:
    text = f.read()

audio = client.text_to_speech.convert(
    text=text,
    voice_id="JBFqnCBsd6RMkjVDRZzb",  # a voice from the docs example, replace it with your Voice ID
    model_id="eleven_v4",
    output_format="mp3_44100_128",
)

with open("$OUTPUT_DIR/voiceover.mp3", "wb") as out:
    for chunk in audio:
        out.write(chunk)
print("Voiceover ready")
PYEOF

# Step 3: FFmpeg builds the video
echo "Assembling the video..."
DURATION=$(ffprobe -v error -show_entries format=duration \
    -of csv=p=0 "$OUTPUT_DIR/voiceover.mp3" | awk '{print int($1+1)}')

ffmpeg \
    -loop 1 -i "$BACKGROUND" \
    -i "$OUTPUT_DIR/voiceover.mp3" \
    -c:v libx264 -tune stillimage \
    -c:a aac -b:a 192k \
    -pix_fmt yuv420p \
    -t "$DURATION" \
    "$OUTPUT_DIR/video.mp4" \
    -y -loglevel error

echo "Done: $OUTPUT_DIR/video.mp4 (${DURATION}s)"

Running it:

bash
chmod +x tts-video-pipeline.sh
./tts-video-pipeline.sh "How RAG works in Claude" background.jpg

A voice journal and a podcast in 5 minutes

You speak your thoughts out loud → Whisper (STT, Speech-to-Text, turning speech into text) transcribes them → Claude edits them for listening → ElevenLabs gives them a professional voiceover.

python
import whisper
import anthropic
from elevenlabs.client import ElevenLabs
import os

# Models
whisper_model = whisper.load_model("small")
claude = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

def voice_to_podcast(audio_file: str, output_file: str = "podcast.mp3"):
    """
    audio_file: a voice recording (WAV/MP3)
    output_file: the finished podcast episode
    """

    # Step 1: transcribe the voice note
    print("Transcribing...")
    result = whisper_model.transcribe(audio_file, language="en")
    raw_text = result["text"]
    print(f"Transcript ({len(raw_text.split())} words): {raw_text[:100]}...")

    # Step 2: Claude edits it for listening
    print("Editing...")
    response = claude.messages.create(
        model="claude-sonnet-5-5",  # current model IDs: see the Anthropic documentation
        max_tokens=2000,
        messages=[{
            "role": "user",
            "content": (
                "This is a transcript of spoken speech. Edit it for a podcast:\n"
                "- Remove filler words (um, uh, like, you know)\n"
                "- Fix broken-off sentences\n"
                "- Keep the conversational tone and the structure of the thought\n"
                "- Don't add anything new, only edit\n\n"
                f"Text:\n{raw_text}"
            )
        }]
    )
    edited_text = "".join(b.text for b in response.content if b.type == "text")

    # Step 3: ElevenLabs gives it a professional voiceover
    print("Recording the voiceover...")
    audio = eleven.text_to_speech.convert(
        text=edited_text,
        voice_id="JBFqnCBsd6RMkjVDRZzb",  # a voice from the docs example, replace it with your Voice ID
        model_id="eleven_v4",
        output_format="mp3_44100_128",
    )
    with open(output_file, "wb") as out:
        for chunk in audio:
            out.write(chunk)

    print(f"Podcast ready: {output_file}")
    return edited_text

# Run it
edited = voice_to_podcast("my-thoughts.wav", "podcast-ep01.mp3")
print("\nEdited text saved to podcast-ep01-script.txt")
with open("podcast-ep01-script.txt", "w", encoding="utf-8") as f:
    f.write(edited)

Services built on TTS

TTS isn't just a tool for yourself. It's the foundation for three kinds of services. There are no prices here on purpose: work out the price of your service with the lesson How to set a price, and nobody can guarantee income.

Option 1: Voicing books and courses

The client brings a manuscript, you deliver audio chapter by chapter. Claude edits the text for listening, ElevenLabs voices it. Your cost is calculated in credits: the number of characters in the manuscript against your plan's credits. You need a paid plan with a commercial license and confirmation that the client holds the rights to the text.

The idea: dozens of hours of a narrator's work are replaced by a few hours of your pipeline. Check the quality by ear: mistakes in stress and in proper names are still your responsibility.

Option 2: Content localization

A creator or a business wants to translate a video course into several languages. You use Eleven v4 with the same voice. A native speaker checks translation and pronunciation, otherwise the mistakes will go out with the publication.

Option 3: TTS SaaS (Software-as-a-Service) for small businesses

Realtors, tutors, small businesses all record announcements and presentations by voice. You can wrap the ElevenLabs API in a simple web interface: upload text → pick a voice → download the MP3. Before launching, check the ElevenLabs API terms on resale and commercial use, and work out your costs with the lesson What a client costs and what they bring in.


Practice

  1. Sign up at elevenlabs.io → get an API key (the Free plan is enough to start)

  2. Install the dependencies:

    bash
    pip install elevenlabs anthropic
  3. Run the basic "Claude writes → ElevenLabs voices" script from the theory section. Listen to the result.

  4. Clone your voice (requires Starter or higher): record 60 seconds of clear speech (read any text), upload it to ElevenLabs in the Voices section (instant clone). Replace voice_id in the code with your Voice ID.

  5. Create the tts-video-pipeline.sh file, make it executable, run it with any topic. Check the resulting MP4.

  6. (Advanced) Install Kokoro TTS locally, voice an English text, and compare it with ElevenLabs. Note the difference in quality.


Tools and resources


Key takeaways

TTS is the last manual step in the content pipeline. ElevenLabs closes it. One voice → many videos, 90+ languages with Eleven v4, no studio.

Voice cloning = a short sample → endless voiceover. Record once, and your voice works without you, as long as you have the voice owner's permission.

Price TTS services (book narration, localization, TTS SaaS) from your costs: plan credits, your time, a native speaker's review. Nobody guarantees prices or income.

A free starter stack: macOS say for tests, Kokoro TTS for offline prototypes in English, ElevenLabs Free (10,000 credits a month, no commercial license) for your first experiments.


Next lesson

→ Music with AI: Suno and sound design: background music, jingles and sound effects for videos and podcasts

The mark stays in this browser only and is never sent anywhere. My progress