The gist
Most support calls are the same questions over and over. "How do I reset my password?" "Where's my order?" "How do I cancel my subscription?" A live agent spends most of their working day on these, burns out, and the company pays for every hour.
A voice AI agent answers these calls 24/7, without burning out, in many languages. You pay for call minutes at the rates of several services at once (more on them below), so the cost has to be calculated for the full stack, not from a single line item. The agent hands complex cases to a live person.
Key concepts
- Vapi: a platform for voice AI agents: inbound calls, responses in a fraction of a second, many languages
- Bland.ai: mass outbound calling (call campaigns, surveys, reminders)
- Transfer on failure: the AI doesn't know the answer → it instantly transfers to an agent
- RAG integration: the AI answers from the company's knowledge base
- Transcription and analytics: every call becomes text → analyzed by topic
- Cost: the sum of five line items (platform, speech recognition, model, voice, telephony) vs an agent's hourly rate
What a minute of a call actually costs (often hidden):
- Platform fee (Vapi, Bland.ai and similar): for Vapi, as of October 2026, it's $0.05 per minute, and that's only the platform's fee
- STT (speech recognition: Whisper, Deepgram and others): separate, at the provider's rate
- LLM (Claude Sonnet, Haiku or another model): depends on the length of the conversation and the knowledge base; model prices: What's current
- TTS (voice synthesis: ElevenLabs and others): separate, at the provider's rate
- Telephony (Twilio, Telnyx and others): the phone number and per-minute charges
The platform is only the first line; the other four are billed separately. Add up all five at your chosen providers' current rates: that's the real price of your minute. The only fair comparison with an agent's hourly rate is for the full stack: on an expensive stack, the savings melt away fast. Calculate your stack before launch, not after.
Theory
Why traditional call centers break down
What usually happens in support (measure the numbers on your own calls; there are no universal figures):
- Some customers hang up if they can't get through in a couple of minutes
- A call with a live agent costs money: shifts, training, turnover
- A large share of calls are repeat questions from the FAQ
- Monotonous work wears agents down, and they make more mistakes
- An AI working from a ready knowledge base answers consistently, but it makes mistakes too: check its answers against call recordings
The problem isn't the people, it's the system. Doing the same thing 50 times a day makes anyone worse at it.
Vapi: inbound voice calls
Vapi is a platform for building voice AI agents. A customer calls your number, and the AI answers.
How it works:
Customer calls → Vapi picks up → STT (speech to text) → LLM (thinks) → TTS (text to speech) → customer hears the answerThe whole loop takes a fraction of a second. With good setup the delay is barely noticeable, but on long answers and bad connections you can hear the pause.
Setting it up through the API:
import requests
import json
VAPI_API_KEY = "your_vapi_key"
def create_support_agent(
company_name: str,
knowledge_base: str,
transfer_phone: str # The live agent's number
) -> dict:
"""Creates a voice support agent through Vapi"""
headers = {
"Authorization": f"Bearer {VAPI_API_KEY}",
"Content-Type": "application/json"
}
# Field and model names in Vapi change: before launching, check the configuration
# against the docs at docs.vapi.ai (model, voice, transcription, call transfer)
agent_config = {
"name": f"{company_name} Support Agent",
"model": {
"provider": "anthropic",
"model": "claude-sonnet-5-5",
"messages": [{"role": "system", "content": f"""You are a customer support agent for {company_name}.
KNOWLEDGE BASE:
{knowledge_base}
RULES:
1. Be friendly and professional
2. Answer only from the knowledge base; don't make things up
3. If you don't know the answer, say so honestly and offer to connect them with an agent
4. Always check that the customer understood
5. If the customer is angry, show empathy and don't argue
TRANSFER TO AN AGENT (say this word for word):
"Let me connect you with one of our specialists who can help with this.
Please hold."
Speak naturally. Pause. Don't rush."""}],
"temperature": 0.3, # Less variation for support
# Transferring the call to a live agent: the transferCall tool
"tools": [{
"type": "transferCall",
"destinations": [{
"type": "number",
"number": transfer_phone,
"description": "Transfer the call to an agent if the AI doesn't know the answer or the customer asks for a person",
"message": "Connecting you with a specialist"
}]
}]
},
"voice": {
"provider": "11labs", # ElevenLabs
"voiceId": "voice_ID_from_ElevenLabs" # pick a voice that sounds natural in your customers' language
},
"transcriber": {
"provider": "deepgram",
"language": "en"
},
"firstMessage": f"Hi! You've reached {company_name} support. How can I help?",
"endCallMessage": "Thanks for calling! If you have any more questions, give us a call. Goodbye!",
"maxDurationSeconds": 600 # 10 minutes max
}
response = requests.post(
"https://api.vapi.ai/assistant",
headers=headers,
json=agent_config
)
return response.json()
def assign_phone_number(assistant_id: str, phone_number: str) -> dict:
"""Links a phone number to the agent"""
headers = {
"Authorization": f"Bearer {VAPI_API_KEY}",
"Content-Type": "application/json"
}
response = requests.post(
"https://api.vapi.ai/phone-number",
headers=headers,
json={
"assistantId": assistant_id,
"number": phone_number,
"provider": "twilio" # or vonage, telnyx
}
)
return response.json()
# Example: an agent for an online store
knowledge_base = """
ORDER QUESTIONS:
- Order status: ask for the order number, go to track.example.com/[number]
- Delivery time: local 1-2 days, nationwide 3-5 days, international 7-14 days
- Changing the address: possible until the order ships, not after
PAYMENT QUESTIONS:
- Methods: Visa/Mastercard, PayPal, Apple Pay, Google Pay
- Returns: within 14 days under our return policy
- Receipt: sent to your email automatically
TECHNICAL PROBLEMS:
- The site isn't working: try clearing your browser cache
- The email didn't arrive: check your Spam folder
- Can't log in: reset your password at example.com/reset
NEEDS AN AGENT:
- Complaints about product quality
- Legal questions
- Situations that require compensation
"""
agent = create_support_agent(
company_name="My Store",
knowledge_base=knowledge_base,
transfer_phone="+15555550123"
)
print(f"Agent created: {agent.get('id')}")
print(f"Name: {agent.get('name')}")Bland.ai: mass outbound calls
Bland.ai specializes in outbound calls. The main uses:
- Reminders: "You have a doctor's appointment tomorrow at 2:00 PM"
- Order confirmations: "Your order #12345 is ready for pickup"
- Collecting feedback: a call after a purchase with 3 questions
- Win-back: "We haven't seen you in a while, here's a personal offer"
import requests
BLAND_API_KEY = "your_bland_key"
def make_batch_calls(
contacts: list[dict], # [{"phone": "+1...", "name": "John", ...}]
call_purpose: str,
script_template: str
) -> dict:
"""Launches a mass calling campaign through Bland.ai"""
# Bland's request schema changes: check the fields and endpoint against the docs at docs.bland.ai
headers = {
"Authorization": BLAND_API_KEY,
"Content-Type": "application/json"
}
# Build a task for each contact
call_objects = []
for contact in contacts:
# Personalize the script for each person
personalized_script = script_template.format(**contact)
call_objects.append({
"phone_number": contact["phone"],
"task": personalized_script
})
# Launch the batch: shared settings go in global, calls go in call_objects
response = requests.post(
"https://api.bland.ai/v2/batches/create",
headers=headers,
json={
"global": {
"task": script_template, # fallback script
"voice": "voice_ID_from_dashboard", # pick a voice in the Bland dashboard
"language": "en", # see the docs for language codes
"max_duration": 2, # 2 minutes max
"record": True # record the call (see the warning below)
},
"call_objects": call_objects,
"description": call_purpose[:60]
}
)
return response.json()
# Example: appointment reminders
contacts_for_reminder = [
{"phone": "+15555550111", "name": "John", "date": "tomorrow", "time": "2:00 PM", "service": "a haircut"},
{"phone": "+15555550122", "name": "Maria", "date": "on Wednesday", "time": "11:30 AM", "service": "a manicure"},
]
reminder_script = """
Hi, {name}! This is Beauty Studio calling.
Just a reminder that you have an appointment for {service} {date} at {time}.
If your plans have changed, please give us a call or send us a text.
See you soon! Goodbye.
"""
result = make_batch_calls(
contacts=contacts_for_reminder,
call_purpose="Appointment reminders — 2026-10-15",
script_template=reminder_script
)
print(f"Batch created: {result.get('batch_id')}")
print(f"Calls scheduled: {len(contacts_for_reminder)}")⚠️ Outbound calling campaigns and call recording are regulated by law: you need the customer's consent, a notice that the call is recorded, and a disclosure that an AI voice is speaking; the rules depend on the country (and in the US, on the state). Before launching, check the lesson AI regulation and compliance and talk to your lawyer.
Integrating a knowledge base (RAG)
A voice AI gets much more capable when it has access to up-to-date company data.
import anthropic
from typing import Optional
# A simple RAG system for a voice agent
def create_rag_support_handler(
documents: list[str], # List of knowledge base documents
) -> callable:
"""Creates a question handler that searches the knowledge base"""
client = anthropic.Anthropic()
# In a real project this would be a vector database (Pinecone, Chroma)
# To keep it simple, we pass all documents into the context
knowledge_context = "\n\n---\n\n".join(documents)
def handle_voice_query(
user_query: str,
conversation_history: list[dict],
customer_info: Optional[dict] = None
) -> dict:
"""Handles a request from the voice agent"""
system_prompt = f"""You are a support agent. Answer ONLY based on the knowledge base.
If the information isn't there, say honestly that you don't know and offer to connect them with an agent.
IMPORTANT FOR VOICE ANSWERS:
- The answer should sound natural when read out loud
- Don't use list markers (•, -, *); they don't sound good
- Write numbers as words: not "14:00" but "at two in the afternoon"
- 3-4 sentences max per answer
KNOWLEDGE BASE:
{knowledge_context}"""
messages = conversation_history + [{
"role": "user",
"content": user_query
}]
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=200, # Voice answers are short
system=system_prompt,
messages=messages
)
answer = "".join(b.text for b in response.content if b.type == "text")
# Detect whether a transfer to an agent is needed
transfer_signals = [
"don't know", "can't help", "need a specialist",
"connect you with an agent", "pass your situation"
]
needs_transfer = any(signal in answer.lower() for signal in transfer_signals)
return {
"response_text": answer,
"needs_human_transfer": needs_transfer,
"transfer_reason": "The AI couldn't answer" if needs_transfer else None
}
return handle_voice_query
# Example usage
faq_documents = [
"""RETURNS
You can return an item within fourteen days of purchase.
The item must be in its original packaging with no signs of use.
We refund the money within three business days to the card you paid with.""",
"""SHIPPING
Local delivery takes one to two days.
Nationwide delivery takes three to seven days.
Shipping is free on orders over fifty dollars.
You can track your order on the website under My Orders.""",
"""PAYMENT
We accept Visa, Mastercard and American Express.
You can also pay with PayPal or Apple Pay.
Installment plans are available through a partner for three, six or twelve months."""
]
handler = create_rag_support_handler(faq_documents)
# Test conversation
conversation = []
queries = [
"I want to return something I bought a week ago",
"what if I don't have the packaging anymore?"
]
for query in queries:
print(f"\nCustomer: {query}")
result = handler(query, conversation)
print(f"AI: {result['response_text']}")
if result['needs_human_transfer']:
print("⚠️ TRANSFER TO AN AGENT")
# Update the conversation history
conversation.append({"role": "user", "content": query})
conversation.append({"role": "assistant", "content": result['response_text']})Call transcription and analytics
Every call turns into data. Claude analyzes thousands of transcripts and finds patterns.
def analyze_call_transcripts(transcripts: list[str]) -> dict:
"""Analyzes a batch of call transcripts and finds patterns"""
client = anthropic.Anthropic()
# Combine the transcripts with separators
combined = "\n\n=== CALL ===\n\n".join(transcripts[:20]) # First 20
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1000,
messages=[{
"role": "user",
"content": f"""Analyze these support call transcripts.
TRANSCRIPTS:
{combined}
I need:
1. **Top 10 questions** asked most often (with percentages)
2. **Top 5 problems** the AI couldn't solve (needed a person)
3. **Average call length** by category
4. **Emotional tone**: how many angry/satisfied/neutral customers
5. **Recommendations**: what to add to the knowledge base so the AI can handle more
6. **Red flags**: what urgently needs fixing in the product/service
Format: a structured report with numbers."""
}]
)
return "".join(b.text for b in response.content if b.type == "text")
# Cost vs ROI
def calculate_roi(
current_calls_per_month: int,
avg_call_duration_minutes: float,
operator_hourly_cost: float,
ai_cost_per_minute: float, # the sum of the five per-minute cost items; calculate it for your stack
ai_deflection_rate: float = 0.75 # an assumption for the example: measure it on your own calls
) -> dict:
"""Calculates the ROI of rolling out voice AI (all numbers in the example are made up)"""
# Current costs
total_minutes = current_calls_per_month * avg_call_duration_minutes
current_cost = (total_minutes / 60) * operator_hourly_cost
# After AI
ai_handled = current_calls_per_month * ai_deflection_rate
human_handled = current_calls_per_month * (1 - ai_deflection_rate)
ai_cost = ai_handled * avg_call_duration_minutes * ai_cost_per_minute
human_cost = (human_handled * avg_call_duration_minutes / 60) * operator_hourly_cost
new_total_cost = ai_cost + human_cost
savings = current_cost - new_total_cost
return {
"current_monthly_cost": round(current_cost, 2),
"new_monthly_cost": round(new_total_cost, 2),
"monthly_savings": round(savings, 2),
"annual_savings": round(savings * 12, 2),
"ai_handled_calls": int(ai_handled),
"human_calls_freed": int(ai_handled),
"roi_percent": round((savings / new_total_cost) * 100, 1)
}
# Example calculation
roi = calculate_roi(
current_calls_per_month=2000,
avg_call_duration_minutes=4.5,
operator_hourly_cost=15, # an assumption for the example: plug in the real rate
ai_cost_per_minute=0.20 # an assumption for the example: plug in the total for your stack
)
print("=== ROI ANALYSIS ===")
for k, v in roi.items():
print(f"{k}: {v}")Practice
Assignment: build a voice support agent for simple questions
What we'll build in 35 minutes:
1. Sign up for Vapi (5 min) 2. Create the agent through the API (10 min) 3. Test it in the browser (10 min) 4. Look at the test call analytics (10 min)
Step 1: Sign up
# Go to vapi.ai
# Create an account
# Get an API key in Dashboard → API Keys
# Save it to .env:
echo "VAPI_API_KEY=your_key_here" >> .envStep 2: Create the agent
# setup_agent.py
import os
import requests
from dotenv import load_dotenv
load_dotenv()
VAPI_API_KEY = os.environ["VAPI_API_KEY"]
# Create the agent
agent = requests.post(
"https://api.vapi.ai/assistant",
headers={"Authorization": f"Bearer {VAPI_API_KEY}"},
json={
"name": "Test support agent",
"model": {
"provider": "anthropic",
"model": "claude-sonnet-5-5",
"messages": [{"role": "system", "content": """You are a support agent for a test store.
You know:
- Shipping: 2-5 days nationwide, free on orders over $50
- Returns: 14 days, item in its packaging
- Payment: cards, PayPal, Apple Pay
- Hours: Mon-Fri 9-6, Sat 10-3
If a question is off-topic, say you're connecting them with an agent.
Keep it short and to the point. Be friendly."""}]
},
"voice": {
"provider": "11labs",
"voiceId": "voice_ID_from_ElevenLabs"
},
"firstMessage": "Hi there! Test support line. How can I help?"
}
).json()
print(f"Agent ID: {agent['id']}")
print(f"To make a test call in the browser:")
print(f"Use Dashboard → Test Phone Call")Step 3: Test it in the Vapi Dashboard
The Vapi dashboard has a test call feature ("Test your assistant" or something similar; button names change). Talk to the agent right in your browser. Ask:
- "How do I return an item?"
- "When will my order arrive?"
- "I want to talk to a real person"
Step 4: Look at the analytics
# get_call_analytics.py
import requests
import os
VAPI_API_KEY = os.environ["VAPI_API_KEY"]
# Get the list of recent calls
calls = requests.get(
"https://api.vapi.ai/call",
headers={"Authorization": f"Bearer {VAPI_API_KEY}"},
params={"limit": 10}
).json()
for call in calls.get("results", []):
print(f"\nCall: {call.get('id')}")
print(f"Duration: {call.get('duration', 0)} sec")
print(f"Status: {call.get('status')}")
if call.get('transcript'):
print(f"Transcript:\n{call['transcript'][:300]}...")Tools and resources
- Vapi: vapi.ai (inbound calls; as of October 2026 the platform charges $0.05 per minute, while speech recognition, the model, voice and telephony are billed separately)
- Bland.ai: bland.ai (outbound calls, pricing on the website)
- Retell AI: retellai.com (an alternative to Vapi)
- ElevenLabs: elevenlabs.io (voice synthesis; see the website for pricing and commercial use terms)
- Deepgram: deepgram.com (speech recognition, pricing on the website)
- OpenAI: Whisper speech recognition and transcription models (pricing on OpenAI's website)
- Twilio: twilio.com (phone numbers and per-minute charges, pricing on the website)
- Anthropic API: the model names in the code are as of October 2026; for current prices and versions, see What's current
Key takeaways
Voice AI isn't a replacement for people, it's a filter. People handle the hard stuff: empathy, unusual situations, sales. AI handles the routine: FAQs, statuses, reminders. Each does its own job better.
Cost is the main argument, but calculate the full stack: platform + speech recognition + model + voice + telephony vs an agent's hourly rate where you are. On a cheap stack the difference can be noticeable; on an expensive one the savings melt away. Don't trust the marketing "$0.05/min": as of October 2026 that's only Vapi's platform fee, without everything else.
The main rule for rolling it out: start with the simplest and most frequent questions. Don't try to automate everything at once. Take the top 5 questions from your FAQ; that's enough for a first agent.
Next lesson
The mark stays in this browser only and is never sent anywhere. My progress