Library · Customer support and voice agents

Voice AI agents: Vapi, Retell, Bland.ai, real phone calls

Builder65 minUpdated: October 2026
63 of 105 in the library

Module: Voice & Real-Time AI | Time: ~25 min theory + 40 min practice


The gist

Chatbots write. Voice agents call.

They're different classes of tools. A chatbot waits for the customer to write. A voice agent can dial the number itself shortly after the customer submits a request and agrees to a call, then qualify them, ask the right questions and book a meeting. No human involved.

In this lesson we build a real voice agent with Vapi: it calls, it talks, and after the call Claude analyzes the transcript and saves the data to your CRM.

🎨 Picture this: a voice agent is an invisible employee. Works 24/7, never gets sick, never gets tired, sticks to the script, and handles many calls at once (the limit depends on the plan). It's paid by the minute, not a fixed salary.


Key concepts

  • Voice AI agent: a system that calls real people and talks with them by voice
  • Latency: the delay before a response. The shorter it is, the more natural the conversation; long pauses sound robotic
  • STT (Speech-to-Text): turns the customer's speech into text
  • TTS (Text-to-Speech): turns the LLM's answer into a voice
  • Vapi: a platform for developers: API-first, Claude plugs in as the model provider, configurable for different scenarios
  • Webhook: the URL Vapi sends events to (the call status, and after the call, a report with the transcript)
  • First message (firstMessage): what the agent says the moment the customer picks up

Theory

Why voice agents are a different story

Chatbots and voice agents solve different problems.

Chatbot: the customer has to come to you, write, and wait for an answer. The initiative is on the customer's side.

Voice agent: the system initiates contact itself, runs the conversation itself, and decides on its own what to ask next. The initiative is on the business's side.

That's why voice agents are used where response speed is critical:

  • Qualifying inbound leads: a request comes in, and the agent calls almost right away (if the customer agreed to a call). A fast response usually helps, but the effect depends on the niche: check it against your own data
  • Appointment reminders: a call a day ahead helps reduce no-shows
  • Post-purchase calls: gathers reviews and ratings without a call center
  • Outbound campaigns: the cost is calculated per minute (see the cost section), and you can only call people who have given consent

🎨 Picture this: a website form without a voice agent is a mailbox. A form with a voice agent is a phone that rings on its own the moment a letter drops in.


The main technical constraint: latency

A voice conversation needs real-time answers. People are comfortable with a short pause in a conversation, about half a second. Anything longer feels like "the robot is thinking."

The latency chain in a voice agent:

Code
Customer speaks
  → STT (speech recognition): hundreds of milliseconds
  → LLM (generating the answer): hundreds of milliseconds
  → TTS (voice synthesis): hundreds of milliseconds
  → Customer hears the answer
  
Total: you need to fit into roughly one second; the faster, the better

The exact delays depend on the providers and model you pick, so measure them on your own stack with test calls.

That's exactly why platforms like Vapi and Retell put most of their effort into optimizing this chain. They don't just wrap Whisper + GPT + ElevenLabs; they optimize every step at the infrastructure level.


Comparing platforms

Platform Focus Pricing (as of October 2026) Models
Vapi.ai Developers, API-first $0.05/min for the platform (as of October 2026); everything else is billed separately at the providers' prices Claude, GPT, Gemini and others (list in the docs)
Retell AI Businesses, quick start per minute, depending on the model and voice you choose (see the website) chosen in the platform settings
Bland.ai Outbound/SDR calls per-minute plans (model, recognition and voice included) and plans with a monthly fee (see the website) the platform's models
ElevenLabs (voice agents) Voice quality see the plans on the website chosen in the settings
Twilio + Claude DIY Maximum control Twilio telephony plus the speech and model providers Any

⚠️ The platform price isn't the full price per minute. Vapi's $0.05/min covers orchestration only; speech recognition, the model, the voice and telephony are billed separately. See the breakdown in the "The real cost of a minute" section. Current prices and versions: What's current.

Recommendation: Vapi for developers: a clear API, Claude plugs in as the model provider, and there's documentation and a community. Retell if you need a quick start without code. Bland.ai if your goal is outbound sales. Platform prices and features change, so check them on the websites.


Vapi: how the platform works

Vapi has three components:

  1. Assistant: the agent's configuration: system prompt, voice, first message, LLM settings
  2. Call: a specific call: who to call, which assistant to use
  3. Webhook: what to do after the call: transcript, status, recording

Everything is managed through the REST API or the Dashboard at vapi.ai. For production, use the API. For testing, the Dashboard.


Practice

Step 1: Sign up and run a first test

  1. Sign up at vapi.ai. New accounts get starter credits (as of October 2026: $5), which is enough for test calls
  2. In the Dashboard → API Keys → copy your key
  3. In the Dashboard → Phone Numbers → buy a number or connect your own Twilio number. Copy the number's ID (phoneNumberId): you need it for outbound calls

Save the key in .env:

bash
VAPI_KEY=your-vapi-key-here
VAPI_PHONE_NUMBER_ID=your-phone-number-id
ANTHROPIC_API_KEY=sk-ant-your-key-here

Step 2: A minimal voice agent

We create an assistant and make the first test call.

python
# voice_agent.py
import requests
import os
from dotenv import load_dotenv

load_dotenv()

VAPI_KEY = os.getenv("VAPI_KEY")
HEADERS = {"Authorization": f"Bearer {VAPI_KEY}"}

def create_lead_qualifier() -> str:
    """
    Creates a voice agent for qualifying leads.
    Returns the ID of the created assistant.
    """
    assistant = requests.post(
        "https://api.vapi.ai/assistant",
        headers=HEADERS,
        json={
            "name": "Lead Qualifier EN",
            "model": {
                "provider": "anthropic",
                "model": "claude-sonnet-5-5",   # check that the model is on Vapi's list; current models: the What's current page
                "messages": [{"role": "system", "content": """You are a professional sales rep.
Your job is to qualify an inbound lead in 3-5 minutes.

Ask these four questions during the conversation:
1. What budget is the customer considering?
2. What's the timeline: when do they plan to decide?
3. What exactly are they looking for: specific requirements?
4. Who makes the final decision?

Conversation rules:
- Keep it short and friendly, don't sound scripted
- Don't ask all the questions in a row: weave them into the conversation
- If the customer already answered a question, don't ask it again
- At the end, offer to book a meeting with a manager
- Keep the call under 5 minutes
- If asked whether you're a person, honestly say you're an AI assistant"""}],
                "temperature": 0.7,
                "maxTokens": 150
            },
            "voice": {
                "provider": "11labs",
                "voiceId": "pNInz6obpgDQGcFmaJgB"
            },
            "firstMessage": "Hi! This is the company's AI assistant. You submitted a request on our website. Do you have a couple of minutes to talk?",
            "endCallMessage": "Great, I've got everything I need. One of our managers will get in touch within the hour. Have a great day!",
            "maxDurationSeconds": 300
        }
    ).json()

    assistant_id = assistant["id"]
    print(f"Assistant created: {assistant_id}")
    return assistant_id


def make_call(assistant_id: str, phone_number: str) -> dict:
    """
    Starts a call to a customer.
    phone_number in E.164 format: +15551234567
    Only call people who have given consent (see the section on laws below).
    """
    call = requests.post(
        "https://api.vapi.ai/call",
        headers=HEADERS,
        json={
            "assistantId": assistant_id,
            "phoneNumberId": os.getenv("VAPI_PHONE_NUMBER_ID"),   # the number we call from
            "customer": {
                "number": phone_number,
                "name": "Customer"  # optional, for the logs
            }
        }
    ).json()

    print(f"Call started: {call['id']}")
    print(f"Status: {call['status']}")
    return call


def get_call_transcript(call_id: str) -> str:
    """
    Gets the transcript of a finished call.
    """
    call = requests.get(
        f"https://api.vapi.ai/call/{call_id}",
        headers=HEADERS
    ).json()

    if call.get("artifact", {}).get("transcript"):
        return call["artifact"]["transcript"]
    return ""


# Run it
if __name__ == "__main__":
    assistant_id = create_lead_qualifier()

    # For testing, call your own number
    call = make_call(assistant_id, "+15551234567")
    print(f"\nCall ID for getting the transcript: {call['id']}")

Installing the dependencies:

bash
pip install requests python-dotenv

Step 3: RAG: the agent knows your product

The agent needs to know your services, prices and FAQ. You do this through a system prompt with context.

python
# knowledge_loader.py
from pathlib import Path

def load_knowledge_base(filename: str) -> str:
    """Reads a knowledge base file"""
    path = Path("knowledge") / filename
    if path.exists():
        return path.read_text(encoding="utf-8")
    return ""


def build_system_prompt() -> str:
    """Builds the system prompt with the knowledge base"""

    services = load_knowledge_base("services.md")
    faq = load_knowledge_base("faq.md")
    objections = load_knowledge_base("objections.md")

    return f"""You are a consultant for a real estate agency in Ecuador.

OUR SERVICES AND PRICES:
{services}

FREQUENTLY ASKED QUESTIONS:
{faq}

HOW TO HANDLE OBJECTIONS:
{objections}

RULES:
- Never quote prices higher than the ones listed; always say "starting at X"
- If the question isn't in the FAQ, say "I'll check the details and we'll call you back"
- Always offer to book a meeting or a video consultation
- Speak English unless the customer switches to another language (for example, Spanish)
- Don't read out lists: weave the information into the conversation"""


# Example: create an assistant with the knowledge base
def create_realty_agent() -> str:
    import requests
    import os

    VAPI_KEY = os.getenv("VAPI_KEY")

    assistant = requests.post(
        "https://api.vapi.ai/assistant",
        headers={"Authorization": f"Bearer {VAPI_KEY}"},
        json={
            "name": "Realty Consultant",
            "model": {
                "provider": "anthropic",
                "model": "claude-sonnet-5-5",
                "messages": [{"role": "system", "content": build_system_prompt()}]
            },
            "voice": {
                "provider": "11labs",
                "voiceId": "pNInz6obpgDQGcFmaJgB"
            },
            "firstMessage": "Hi! This is the agency's AI consultant. How can I help?"
        }
    ).json()

    return assistant["id"]

Create a knowledge/ folder with three files:

Code
knowledge/
  services.md    — list of services with prices
  faq.md         — 10-15 frequent questions with answers
  objections.md  — typical objections and how to respond

Step 4: Webhook: what to do after the call

After every call, Vapi sends the data to your webhook. This is where Claude analyzes the transcript and saves the result.

python
# webhook_handler.py
from flask import Flask, request, jsonify
import anthropic
import json
import os
from dotenv import load_dotenv

load_dotenv()

app = Flask(__name__)
claude = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))


def analyze_call(transcript: str) -> dict:
    """
    Claude analyzes the transcript and extracts structured data.
    """
    response = "".join(b.text for b in claude.messages.create(
        model="claude-sonnet-5-5",   # current models: the What's current page
        max_tokens=600,
        messages=[{
            "role": "user",
            "content": f"""Analyze this conversation with a customer and return JSON.

TRANSCRIPT:
{transcript}

Return ONLY valid JSON with no explanation:
{{
  "qualified": true or false,
  "budget": "amount or range, or 'not specified'",
  "timeline": "when they plan to decide",
  "requirements": "what exactly they're looking for (2-3 sentences)",
  "decision_maker": true or false,
  "next_action": "what to do next (for example: have a manager call, send a presentation, remove from the funnel)",
  "sentiment": "positive, neutral or negative",
  "summary": "2-3 sentences about the conversation"
}}"""
        }]
    ).content if b.type == "text")

    # Strip markdown if the model added it
    if "```json" in response:
        response = response.split("```json")[1].split("```")[0]
    elif "```" in response:
        response = response.split("```")[1].split("```")[0]

    try:
        return json.loads(response.strip())
    except json.JSONDecodeError:
        return {
            "qualified": False,
            "summary": "Couldn't analyze the transcript",
            "raw_response": response
        }


def save_to_crm(analysis: dict, customer: dict, call_id: str):
    """
    Saves the result to your system.
    Replace this with your real CRM: Notion, Airtable, Google Sheets, HubSpot.
    """
    record = {
        "call_id": call_id,
        "customer_phone": customer.get("number", "unknown"),
        "customer_name": customer.get("name", "unknown"),
        **analysis
    }

    # The simple version: save to a JSON file
    import datetime
    filename = f"calls/call-{call_id}-{datetime.date.today()}.json"
    os.makedirs("calls", exist_ok=True)

    with open(filename, "w", encoding="utf-8") as f:
        json.dump(record, f, ensure_ascii=False, indent=2)

    print(f"Saved: {filename}")

    # If qualified, notify a manager
    if analysis.get("qualified"):
        notify_manager(record)


def notify_manager(record: dict):
    """
    Notifies a manager about a qualified lead.
    Replace with a real channel: Slack, SMS, email.
    """
    # Example: print to the console (in production, a Slack webhook or email)
    print(f"""
=== QUALIFIED LEAD ===
Phone: {record['customer_phone']}
Budget: {record.get('budget', 'not specified')}
Timeline: {record.get('timeline', 'not specified')}
Next step: {record.get('next_action', '')}
Summary: {record.get('summary', '')}
==============================
    """)


@app.post("/vapi-webhook")
def handle_vapi_event():
    """The main handler for events from Vapi"""
    # Vapi puts the event inside the "message" key
    message = (request.json or {}).get("message", {})
    event_type = message.get("type")

    if event_type == "end-of-call-report":
        call = message.get("call", {})
        call_id = call.get("id", "unknown")
        customer = call.get("customer", {})
        transcript = message.get("artifact", {}).get("transcript", "")

        print(f"Call ended: {call_id}")
        print(f"Reason it ended: {message.get('endedReason', 'unknown')}")

        if transcript:
            print("Analyzing the transcript...")
            analysis = analyze_call(transcript)
            save_to_crm(analysis, customer, call_id)
        else:
            print("Empty transcript: the call didn't connect or the customer didn't answer")

    elif event_type == "status-update":
        call_id = message.get("call", {}).get("id", "unknown")
        print(f"Call {call_id} status: {message.get('status', '?')}")

    return jsonify({"status": "ok"})


if __name__ == "__main__":
    # In production, protect the webhook with a secret in the header and don't enable debug
    app.run(port=5000, debug=False)

Installation:

bash
pip install flask anthropic requests python-dotenv

Running it locally (for testing):

bash
# Start the webhook server
python webhook_handler.py

# In another terminal, forward the port through ngrok for testing
ngrok http 5000
# ngrok will give you a public URL like https://abc123.ngrok.io

Register the webhook in Vapi:

python
# Update the assistant so it sends events to your webhook
import requests

VAPI_KEY = "your-key"
ASSISTANT_ID = "your-assistant-id"
WEBHOOK_URL = "https://abc123.ngrok.io/vapi-webhook"

requests.patch(
    f"https://api.vapi.ai/assistant/{ASSISTANT_ID}",
    headers={"Authorization": f"Bearer {VAPI_KEY}"},
    json={"server": {"url": WEBHOOK_URL}}   # older examples use a serverUrl field
)

Step 5: Calling leads automatically

A real scenario: a new request comes in from the website → the agent calls automatically. An automated call is only OK if the person gave consent to be called in their request (see the section on laws below).

python
# auto_caller.py
"""
A script for automated calling.
Connect it to a webhook from your CRM or contact form.
"""
import requests
import os
from dotenv import load_dotenv

load_dotenv()

VAPI_KEY = os.getenv("VAPI_KEY")
ASSISTANT_ID = os.getenv("VAPI_ASSISTANT_ID")
PHONE_NUMBER_ID = os.getenv("VAPI_PHONE_NUMBER_ID")


def call_new_lead(phone: str, name: str = "", source: str = "") -> str:
    """
    Calls a new lead right after they submit a request.
    Returns a call_id for tracking.
    """
    call = requests.post(
        "https://api.vapi.ai/call",
        headers={"Authorization": f"Bearer {VAPI_KEY}"},
        json={
            "assistantId": ASSISTANT_ID,
            "phoneNumberId": PHONE_NUMBER_ID,
            "customer": {
                "number": phone,
                "name": name
            },
            # Pass context to the agent through overrides
            "assistantOverrides": {
                "variableValues": {
                    "lead_source": source,
                    "lead_name": name
                }
            }
        }
    ).json()

    return call.get("id", "")


def batch_call(leads: list) -> list:
    """
    Calls a list of leads.
    leads = [{"phone": "+1...", "name": "...", "source": "..."}, ...]
    """
    results = []

    for lead in leads:
        print(f"Calling: {lead['name']} ({lead['phone']})")
        call_id = call_new_lead(
            phone=lead["phone"],
            name=lead.get("name", ""),
            source=lead.get("source", "")
        )
        results.append({
            "lead": lead,
            "call_id": call_id
        })

        # Pause between calls so we don't overload anything
        import time
        time.sleep(2)

    return results


# Example: calling leads after exporting them from the CRM
if __name__ == "__main__":
    today_leads = [
        {"phone": "+15551111111", "name": "John Smith", "source": "Website"},
        {"phone": "+15552222222", "name": "Maria Garcia", "source": "Ads"},
    ]

    results = batch_call(today_leads)
    print(f"\nCalls started: {len(results)}")
    for r in results:
        print(f"  {r['lead']['name']}: call_id = {r['call_id']}")

The real cost of a minute: breaking down the full stack

Vapi quotes "$0.05/min," but that's only the orchestration layer. The real stack is billed in parts. Each part's price depends on the provider you choose and changes over time, so take the current numbers from the Vapi pricing page and from each provider's website:

Component Cost What it is
Vapi platform $0.05/min Orchestrating the STT→LLM→TTS pipeline
STT (Deepgram) per minute, see the provider's prices Speech recognition
LLM depends on the model The agent's "brain"; for Claude, calculate from the token prices of the model you pick
TTS (ElevenLabs) per minute or by subscription Voice synthesis
Telephony (Twilio) per minute, depends on the country and the number The phone line
TOTAL the sum of all the lines above The real price per minute; it depends heavily on the model and voice

Model prices (as of October 2026), per million tokens, input/output:

  • Claude Haiku 4.5: $1 / $5, for FAQs and simple intents (check that the model is still available: Anthropic says it will be retired "no earlier than 10/15/2026")
  • Claude Sonnet 5.5: $2 / $10, a balance of price and quality for production
  • Claude Opus 5.5: $4 / $20, for complex conversations, more expensive

For the ElevenLabs voice, you can choose a subscription instead of pay-as-you-go: see the Starter and Creator plans at elevenlabs.io/pricing. Current prices and versions: What's current.


Use cases

The size of the effect depends on the niche, the script and the quality of the leads: measure it on your own data instead of taking it from advertising promises.

Lead qualification: A request comes in → Vapi calls shortly after → in a few minutes the agent gathers the budget, timeline and requirements → the manager gets a ready-made lead profile. Result: the manager sees right away who's worth talking to.

Appointment reminders: 24 hours before the appointment, the agent calls, confirms the time and asks whether anything needs to be prepared. Result: some no-shows can be prevented; your own measurements will show how many.

Post-sale follow-up: 7 days after the purchase, the agent calls, checks that everything's fine and collects feedback. Result: feedback without a call center, and problems caught early.

Outbound campaigns: Calculate the cost of 1,000 minutes of conversation using the full stack (Vapi + STT + LLM + TTS + telephony), not the platform price: the platform fee is only part of the total. The agent handles many calls in parallel (the limit depends on the plan); a person would need days for the same number of calls.


Calling laws and ethics

Automated calls are regulated by law, and the rules depend on the country. In the US, for example, there's the TCPA; in Canada, the CRTC rules; in the EU, GDPR and ePrivacy. General principles:

  • only call people who have given consent to be called (for example, they submitted a request and agreed to be contacted);
  • at the start of the conversation, say that this is an AI assistant, and give the person a way to end the conversation and opt out of future calls;
  • if the conversation is being recorded, say so: in a number of countries and US states this is required;
  • treat transcripts and phone numbers as personal data: limit access, and don't send anything extra to the CRM.

This is a general guide, not legal advice: before launching, check the rules in your country and your customers' countries with a lawyer.

Making money with voice agents

You can offer this as a service, but prices and demand depend on the market and the niche, and income isn't guaranteed.

Model 1: One-time setup: the assistant, the knowledge base, a webhook with CRM integration, and test calls before launch.

Model 2: Support and optimization: monitoring transcripts, improving prompts, adding new scenarios.

How to calculate the value for a client: the manager's hours on calls × their hourly rate, versus the cost of conversation minutes on the full stack plus your fee. Calculate honestly using the client's data, and don't promise savings you can't back up. For how to put together a package and set a price, see the Packaging your offer and Pricing and monetization lessons.


Tools and resources

  • Vapi.ai: the platform and API documentation
  • Retell AI: an alternative focused on a quick start
  • Bland.ai: specializes in outbound/SDR calls
  • ElevenLabs: top-quality TTS voices (integrates with Vapi)
  • ngrok: a local tunnel for testing webhooks
  • Flask: a minimal web server for the webhook

Key takeaways

Voice agents aren't upgraded chatbots. They're a different class of tool: they call on their own, run the conversation and pass the data to your CRM. An honest agent tells the person up front that it's an AI and only calls people who have given consent.

Vapi + Claude is a working combination: Claude plugs in as the model provider. If you need conversations in other languages (for example, Spanish), test them with calls on your own scripts.

After the call, Claude again: it analyzes the transcript, qualifies the lead and decides the next step. The voice agent gathers the data, and Claude processes it.

Latency is critical. If your DIY Twilio + Claude + ElevenLabs stack has a delay of a second or more, the conversation sounds unnatural. Use ready-made platforms (Vapi, Retell) that have optimized this at the infrastructure level.


Next lesson

→ Realtime AI: real-time conversations over WebSocket without intermediary services

The mark stays in this browser only and is never sent anywhere. My progress