The gist
Chatbots write. Voice agents call.
They're different classes of tools. A chatbot waits for the customer to write. A voice agent can dial the number itself shortly after the customer submits a request and agrees to a call, then qualify them, ask the right questions and book a meeting. No human involved.
In this lesson we build a real voice agent with Vapi: it calls, it talks, and after the call Claude analyzes the transcript and saves the data to your CRM.
Key concepts
- Voice AI agent: a system that calls real people and talks with them by voice
- Latency: the delay before a response. The shorter it is, the more natural the conversation; long pauses sound robotic
- STT (Speech-to-Text): turns the customer's speech into text
- TTS (Text-to-Speech): turns the LLM's answer into a voice
- Vapi: a platform for developers: API-first, Claude plugs in as the model provider, configurable for different scenarios
- Webhook: the URL Vapi sends events to (the call status, and after the call, a report with the transcript)
- First message (firstMessage): what the agent says the moment the customer picks up
Theory
Why voice agents are a different story
Chatbots and voice agents solve different problems.
Chatbot: the customer has to come to you, write, and wait for an answer. The initiative is on the customer's side.
Voice agent: the system initiates contact itself, runs the conversation itself, and decides on its own what to ask next. The initiative is on the business's side.
That's why voice agents are used where response speed is critical:
- Qualifying inbound leads: a request comes in, and the agent calls almost right away (if the customer agreed to a call). A fast response usually helps, but the effect depends on the niche: check it against your own data
- Appointment reminders: a call a day ahead helps reduce no-shows
- Post-purchase calls: gathers reviews and ratings without a call center
- Outbound campaigns: the cost is calculated per minute (see the cost section), and you can only call people who have given consent
The main technical constraint: latency
A voice conversation needs real-time answers. People are comfortable with a short pause in a conversation, about half a second. Anything longer feels like "the robot is thinking."
The latency chain in a voice agent:
Customer speaks
→ STT (speech recognition): hundreds of milliseconds
→ LLM (generating the answer): hundreds of milliseconds
→ TTS (voice synthesis): hundreds of milliseconds
→ Customer hears the answer
Total: you need to fit into roughly one second; the faster, the betterThe exact delays depend on the providers and model you pick, so measure them on your own stack with test calls.
That's exactly why platforms like Vapi and Retell put most of their effort into optimizing this chain. They don't just wrap Whisper + GPT + ElevenLabs; they optimize every step at the infrastructure level.
Comparing platforms
| Platform | Focus | Pricing (as of October 2026) | Models |
|---|---|---|---|
| Vapi.ai | Developers, API-first | $0.05/min for the platform (as of October 2026); everything else is billed separately at the providers' prices | Claude, GPT, Gemini and others (list in the docs) |
| Retell AI | Businesses, quick start | per minute, depending on the model and voice you choose (see the website) | chosen in the platform settings |
| Bland.ai | Outbound/SDR calls | per-minute plans (model, recognition and voice included) and plans with a monthly fee (see the website) | the platform's models |
| ElevenLabs (voice agents) | Voice quality | see the plans on the website | chosen in the settings |
| Twilio + Claude DIY | Maximum control | Twilio telephony plus the speech and model providers | Any |
⚠️ The platform price isn't the full price per minute. Vapi's $0.05/min covers orchestration only; speech recognition, the model, the voice and telephony are billed separately. See the breakdown in the "The real cost of a minute" section. Current prices and versions: What's current.
Recommendation: Vapi for developers: a clear API, Claude plugs in as the model provider, and there's documentation and a community. Retell if you need a quick start without code. Bland.ai if your goal is outbound sales. Platform prices and features change, so check them on the websites.
Vapi: how the platform works
Vapi has three components:
- Assistant: the agent's configuration: system prompt, voice, first message, LLM settings
- Call: a specific call: who to call, which assistant to use
- Webhook: what to do after the call: transcript, status, recording
Everything is managed through the REST API or the Dashboard at vapi.ai. For production, use the API. For testing, the Dashboard.
Practice
Step 1: Sign up and run a first test
- Sign up at vapi.ai. New accounts get starter credits (as of October 2026: $5), which is enough for test calls
- In the Dashboard → API Keys → copy your key
- In the Dashboard → Phone Numbers → buy a number or connect your own Twilio number. Copy the number's ID (
phoneNumberId): you need it for outbound calls
Save the key in .env:
VAPI_KEY=your-vapi-key-here
VAPI_PHONE_NUMBER_ID=your-phone-number-id
ANTHROPIC_API_KEY=sk-ant-your-key-hereStep 2: A minimal voice agent
We create an assistant and make the first test call.
# voice_agent.py
import requests
import os
from dotenv import load_dotenv
load_dotenv()
VAPI_KEY = os.getenv("VAPI_KEY")
HEADERS = {"Authorization": f"Bearer {VAPI_KEY}"}
def create_lead_qualifier() -> str:
"""
Creates a voice agent for qualifying leads.
Returns the ID of the created assistant.
"""
assistant = requests.post(
"https://api.vapi.ai/assistant",
headers=HEADERS,
json={
"name": "Lead Qualifier EN",
"model": {
"provider": "anthropic",
"model": "claude-sonnet-5-5", # check that the model is on Vapi's list; current models: the What's current page
"messages": [{"role": "system", "content": """You are a professional sales rep.
Your job is to qualify an inbound lead in 3-5 minutes.
Ask these four questions during the conversation:
1. What budget is the customer considering?
2. What's the timeline: when do they plan to decide?
3. What exactly are they looking for: specific requirements?
4. Who makes the final decision?
Conversation rules:
- Keep it short and friendly, don't sound scripted
- Don't ask all the questions in a row: weave them into the conversation
- If the customer already answered a question, don't ask it again
- At the end, offer to book a meeting with a manager
- Keep the call under 5 minutes
- If asked whether you're a person, honestly say you're an AI assistant"""}],
"temperature": 0.7,
"maxTokens": 150
},
"voice": {
"provider": "11labs",
"voiceId": "pNInz6obpgDQGcFmaJgB"
},
"firstMessage": "Hi! This is the company's AI assistant. You submitted a request on our website. Do you have a couple of minutes to talk?",
"endCallMessage": "Great, I've got everything I need. One of our managers will get in touch within the hour. Have a great day!",
"maxDurationSeconds": 300
}
).json()
assistant_id = assistant["id"]
print(f"Assistant created: {assistant_id}")
return assistant_id
def make_call(assistant_id: str, phone_number: str) -> dict:
"""
Starts a call to a customer.
phone_number in E.164 format: +15551234567
Only call people who have given consent (see the section on laws below).
"""
call = requests.post(
"https://api.vapi.ai/call",
headers=HEADERS,
json={
"assistantId": assistant_id,
"phoneNumberId": os.getenv("VAPI_PHONE_NUMBER_ID"), # the number we call from
"customer": {
"number": phone_number,
"name": "Customer" # optional, for the logs
}
}
).json()
print(f"Call started: {call['id']}")
print(f"Status: {call['status']}")
return call
def get_call_transcript(call_id: str) -> str:
"""
Gets the transcript of a finished call.
"""
call = requests.get(
f"https://api.vapi.ai/call/{call_id}",
headers=HEADERS
).json()
if call.get("artifact", {}).get("transcript"):
return call["artifact"]["transcript"]
return ""
# Run it
if __name__ == "__main__":
assistant_id = create_lead_qualifier()
# For testing, call your own number
call = make_call(assistant_id, "+15551234567")
print(f"\nCall ID for getting the transcript: {call['id']}")Installing the dependencies:
pip install requests python-dotenvStep 3: RAG: the agent knows your product
The agent needs to know your services, prices and FAQ. You do this through a system prompt with context.
# knowledge_loader.py
from pathlib import Path
def load_knowledge_base(filename: str) -> str:
"""Reads a knowledge base file"""
path = Path("knowledge") / filename
if path.exists():
return path.read_text(encoding="utf-8")
return ""
def build_system_prompt() -> str:
"""Builds the system prompt with the knowledge base"""
services = load_knowledge_base("services.md")
faq = load_knowledge_base("faq.md")
objections = load_knowledge_base("objections.md")
return f"""You are a consultant for a real estate agency in Ecuador.
OUR SERVICES AND PRICES:
{services}
FREQUENTLY ASKED QUESTIONS:
{faq}
HOW TO HANDLE OBJECTIONS:
{objections}
RULES:
- Never quote prices higher than the ones listed; always say "starting at X"
- If the question isn't in the FAQ, say "I'll check the details and we'll call you back"
- Always offer to book a meeting or a video consultation
- Speak English unless the customer switches to another language (for example, Spanish)
- Don't read out lists: weave the information into the conversation"""
# Example: create an assistant with the knowledge base
def create_realty_agent() -> str:
import requests
import os
VAPI_KEY = os.getenv("VAPI_KEY")
assistant = requests.post(
"https://api.vapi.ai/assistant",
headers={"Authorization": f"Bearer {VAPI_KEY}"},
json={
"name": "Realty Consultant",
"model": {
"provider": "anthropic",
"model": "claude-sonnet-5-5",
"messages": [{"role": "system", "content": build_system_prompt()}]
},
"voice": {
"provider": "11labs",
"voiceId": "pNInz6obpgDQGcFmaJgB"
},
"firstMessage": "Hi! This is the agency's AI consultant. How can I help?"
}
).json()
return assistant["id"]Create a knowledge/ folder with three files:
knowledge/
services.md — list of services with prices
faq.md — 10-15 frequent questions with answers
objections.md — typical objections and how to respondStep 4: Webhook: what to do after the call
After every call, Vapi sends the data to your webhook. This is where Claude analyzes the transcript and saves the result.
# webhook_handler.py
from flask import Flask, request, jsonify
import anthropic
import json
import os
from dotenv import load_dotenv
load_dotenv()
app = Flask(__name__)
claude = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
def analyze_call(transcript: str) -> dict:
"""
Claude analyzes the transcript and extracts structured data.
"""
response = "".join(b.text for b in claude.messages.create(
model="claude-sonnet-5-5", # current models: the What's current page
max_tokens=600,
messages=[{
"role": "user",
"content": f"""Analyze this conversation with a customer and return JSON.
TRANSCRIPT:
{transcript}
Return ONLY valid JSON with no explanation:
{{
"qualified": true or false,
"budget": "amount or range, or 'not specified'",
"timeline": "when they plan to decide",
"requirements": "what exactly they're looking for (2-3 sentences)",
"decision_maker": true or false,
"next_action": "what to do next (for example: have a manager call, send a presentation, remove from the funnel)",
"sentiment": "positive, neutral or negative",
"summary": "2-3 sentences about the conversation"
}}"""
}]
).content if b.type == "text")
# Strip markdown if the model added it
if "```json" in response:
response = response.split("```json")[1].split("```")[0]
elif "```" in response:
response = response.split("```")[1].split("```")[0]
try:
return json.loads(response.strip())
except json.JSONDecodeError:
return {
"qualified": False,
"summary": "Couldn't analyze the transcript",
"raw_response": response
}
def save_to_crm(analysis: dict, customer: dict, call_id: str):
"""
Saves the result to your system.
Replace this with your real CRM: Notion, Airtable, Google Sheets, HubSpot.
"""
record = {
"call_id": call_id,
"customer_phone": customer.get("number", "unknown"),
"customer_name": customer.get("name", "unknown"),
**analysis
}
# The simple version: save to a JSON file
import datetime
filename = f"calls/call-{call_id}-{datetime.date.today()}.json"
os.makedirs("calls", exist_ok=True)
with open(filename, "w", encoding="utf-8") as f:
json.dump(record, f, ensure_ascii=False, indent=2)
print(f"Saved: {filename}")
# If qualified, notify a manager
if analysis.get("qualified"):
notify_manager(record)
def notify_manager(record: dict):
"""
Notifies a manager about a qualified lead.
Replace with a real channel: Slack, SMS, email.
"""
# Example: print to the console (in production, a Slack webhook or email)
print(f"""
=== QUALIFIED LEAD ===
Phone: {record['customer_phone']}
Budget: {record.get('budget', 'not specified')}
Timeline: {record.get('timeline', 'not specified')}
Next step: {record.get('next_action', '')}
Summary: {record.get('summary', '')}
==============================
""")
@app.post("/vapi-webhook")
def handle_vapi_event():
"""The main handler for events from Vapi"""
# Vapi puts the event inside the "message" key
message = (request.json or {}).get("message", {})
event_type = message.get("type")
if event_type == "end-of-call-report":
call = message.get("call", {})
call_id = call.get("id", "unknown")
customer = call.get("customer", {})
transcript = message.get("artifact", {}).get("transcript", "")
print(f"Call ended: {call_id}")
print(f"Reason it ended: {message.get('endedReason', 'unknown')}")
if transcript:
print("Analyzing the transcript...")
analysis = analyze_call(transcript)
save_to_crm(analysis, customer, call_id)
else:
print("Empty transcript: the call didn't connect or the customer didn't answer")
elif event_type == "status-update":
call_id = message.get("call", {}).get("id", "unknown")
print(f"Call {call_id} status: {message.get('status', '?')}")
return jsonify({"status": "ok"})
if __name__ == "__main__":
# In production, protect the webhook with a secret in the header and don't enable debug
app.run(port=5000, debug=False)Installation:
pip install flask anthropic requests python-dotenvRunning it locally (for testing):
# Start the webhook server
python webhook_handler.py
# In another terminal, forward the port through ngrok for testing
ngrok http 5000
# ngrok will give you a public URL like https://abc123.ngrok.ioRegister the webhook in Vapi:
# Update the assistant so it sends events to your webhook
import requests
VAPI_KEY = "your-key"
ASSISTANT_ID = "your-assistant-id"
WEBHOOK_URL = "https://abc123.ngrok.io/vapi-webhook"
requests.patch(
f"https://api.vapi.ai/assistant/{ASSISTANT_ID}",
headers={"Authorization": f"Bearer {VAPI_KEY}"},
json={"server": {"url": WEBHOOK_URL}} # older examples use a serverUrl field
)Step 5: Calling leads automatically
A real scenario: a new request comes in from the website → the agent calls automatically. An automated call is only OK if the person gave consent to be called in their request (see the section on laws below).
# auto_caller.py
"""
A script for automated calling.
Connect it to a webhook from your CRM or contact form.
"""
import requests
import os
from dotenv import load_dotenv
load_dotenv()
VAPI_KEY = os.getenv("VAPI_KEY")
ASSISTANT_ID = os.getenv("VAPI_ASSISTANT_ID")
PHONE_NUMBER_ID = os.getenv("VAPI_PHONE_NUMBER_ID")
def call_new_lead(phone: str, name: str = "", source: str = "") -> str:
"""
Calls a new lead right after they submit a request.
Returns a call_id for tracking.
"""
call = requests.post(
"https://api.vapi.ai/call",
headers={"Authorization": f"Bearer {VAPI_KEY}"},
json={
"assistantId": ASSISTANT_ID,
"phoneNumberId": PHONE_NUMBER_ID,
"customer": {
"number": phone,
"name": name
},
# Pass context to the agent through overrides
"assistantOverrides": {
"variableValues": {
"lead_source": source,
"lead_name": name
}
}
}
).json()
return call.get("id", "")
def batch_call(leads: list) -> list:
"""
Calls a list of leads.
leads = [{"phone": "+1...", "name": "...", "source": "..."}, ...]
"""
results = []
for lead in leads:
print(f"Calling: {lead['name']} ({lead['phone']})")
call_id = call_new_lead(
phone=lead["phone"],
name=lead.get("name", ""),
source=lead.get("source", "")
)
results.append({
"lead": lead,
"call_id": call_id
})
# Pause between calls so we don't overload anything
import time
time.sleep(2)
return results
# Example: calling leads after exporting them from the CRM
if __name__ == "__main__":
today_leads = [
{"phone": "+15551111111", "name": "John Smith", "source": "Website"},
{"phone": "+15552222222", "name": "Maria Garcia", "source": "Ads"},
]
results = batch_call(today_leads)
print(f"\nCalls started: {len(results)}")
for r in results:
print(f" {r['lead']['name']}: call_id = {r['call_id']}")The real cost of a minute: breaking down the full stack
Vapi quotes "$0.05/min," but that's only the orchestration layer. The real stack is billed in parts. Each part's price depends on the provider you choose and changes over time, so take the current numbers from the Vapi pricing page and from each provider's website:
| Component | Cost | What it is |
|---|---|---|
| Vapi platform | $0.05/min | Orchestrating the STT→LLM→TTS pipeline |
| STT (Deepgram) | per minute, see the provider's prices | Speech recognition |
| LLM | depends on the model | The agent's "brain"; for Claude, calculate from the token prices of the model you pick |
| TTS (ElevenLabs) | per minute or by subscription | Voice synthesis |
| Telephony (Twilio) | per minute, depends on the country and the number | The phone line |
| TOTAL | the sum of all the lines above | The real price per minute; it depends heavily on the model and voice |
Model prices (as of October 2026), per million tokens, input/output:
- Claude Haiku 4.5: $1 / $5, for FAQs and simple intents (check that the model is still available: Anthropic says it will be retired "no earlier than 10/15/2026")
- Claude Sonnet 5.5: $2 / $10, a balance of price and quality for production
- Claude Opus 5.5: $4 / $20, for complex conversations, more expensive
For the ElevenLabs voice, you can choose a subscription instead of pay-as-you-go: see the Starter and Creator plans at elevenlabs.io/pricing. Current prices and versions: What's current.
Use cases
The size of the effect depends on the niche, the script and the quality of the leads: measure it on your own data instead of taking it from advertising promises.
Lead qualification: A request comes in → Vapi calls shortly after → in a few minutes the agent gathers the budget, timeline and requirements → the manager gets a ready-made lead profile. Result: the manager sees right away who's worth talking to.
Appointment reminders: 24 hours before the appointment, the agent calls, confirms the time and asks whether anything needs to be prepared. Result: some no-shows can be prevented; your own measurements will show how many.
Post-sale follow-up: 7 days after the purchase, the agent calls, checks that everything's fine and collects feedback. Result: feedback without a call center, and problems caught early.
Outbound campaigns: Calculate the cost of 1,000 minutes of conversation using the full stack (Vapi + STT + LLM + TTS + telephony), not the platform price: the platform fee is only part of the total. The agent handles many calls in parallel (the limit depends on the plan); a person would need days for the same number of calls.
Calling laws and ethics
Automated calls are regulated by law, and the rules depend on the country. In the US, for example, there's the TCPA; in Canada, the CRTC rules; in the EU, GDPR and ePrivacy. General principles:
- only call people who have given consent to be called (for example, they submitted a request and agreed to be contacted);
- at the start of the conversation, say that this is an AI assistant, and give the person a way to end the conversation and opt out of future calls;
- if the conversation is being recorded, say so: in a number of countries and US states this is required;
- treat transcripts and phone numbers as personal data: limit access, and don't send anything extra to the CRM.
This is a general guide, not legal advice: before launching, check the rules in your country and your customers' countries with a lawyer.
Making money with voice agents
You can offer this as a service, but prices and demand depend on the market and the niche, and income isn't guaranteed.
Model 1: One-time setup: the assistant, the knowledge base, a webhook with CRM integration, and test calls before launch.
Model 2: Support and optimization: monitoring transcripts, improving prompts, adding new scenarios.
How to calculate the value for a client: the manager's hours on calls × their hourly rate, versus the cost of conversation minutes on the full stack plus your fee. Calculate honestly using the client's data, and don't promise savings you can't back up. For how to put together a package and set a price, see the Packaging your offer and Pricing and monetization lessons.
Tools and resources
- Vapi.ai: the platform and API documentation
- Retell AI: an alternative focused on a quick start
- Bland.ai: specializes in outbound/SDR calls
- ElevenLabs: top-quality TTS voices (integrates with Vapi)
- ngrok: a local tunnel for testing webhooks
- Flask: a minimal web server for the webhook
Key takeaways
Voice agents aren't upgraded chatbots. They're a different class of tool: they call on their own, run the conversation and pass the data to your CRM. An honest agent tells the person up front that it's an AI and only calls people who have given consent.
Vapi + Claude is a working combination: Claude plugs in as the model provider. If you need conversations in other languages (for example, Spanish), test them with calls on your own scripts.
After the call, Claude again: it analyzes the transcript, qualifies the lead and decides the next step. The voice agent gathers the data, and Claude processes it.
Latency is critical. If your DIY Twilio + Claude + ElevenLabs stack has a delay of a second or more, the conversation sounds unnatural. Use ready-made platforms (Vapi, Retell) that have optimized this at the infrastructure level.
Next lesson
→ Realtime AI: real-time conversations over WebSocket without intermediary services
The mark stays in this browser only and is never sent anywhere. My progress