The gist
Most business owners build chatbots like this: the bot answers FAQs and that's it. That's an airport with one runway. In this lesson you'll learn to build a system with smart routing: the bot understands the intent, pulls out the key details, and knows when it can answer on its own and when to hand off to a real person immediately.
Key concepts
- Intent classification: automatically figuring out what the user wants (a question, a complaint, a purchase, a request for a human) using Claude Haiku or rules
- Entity extraction: pulling specific details out of a message: names, amounts, dates, product SKUs
- Intent routing: the logic that steers the conversation: rules (keywords) vs. AI classification vs. a hybrid approach
- Escalation triggers: the set of conditions under which the bot hands the conversation to a human: customer anger, a direct request, a complex question, a large amount
- State management: keeping the conversation's context between messages with Redis or Supabase
- Multi-channel unification: one routing logic for Telegram, WhatsApp and a website widget
- Fallback handling: what to do when the bot doesn't understand the intent: ask a clarifying question, offer options, or escalate right away
- Quality metrics: escalation rate (% handed to agents), resolution rate (% resolved by the bot), CSAT (the customer's rating)
Theory
What intent classification is and why you need it
When a customer writes "I want to return an item," it's not just text. It's an intent (a return) that calls for a specific response flow. When they write "are you people serious?!", that's an alarm: the customer is irritated, it's risky for the bot to answer, and a real person is needed.
Intent classification means automatically assigning an intent label to every incoming message. Without it, the bot works like a cashier on their first day: it doesn't understand what's being asked and tries to answer everything with the same template.
The main intent categories for a business bot:
- FAQ: questions about prices, shipping terms, business hours
- Complaint: a complaint, a problem, negativity
- Purchase: interest in buying, a request for advice
- Support: a technical question, a problem with an order
- Human request: an explicit request to talk to an agent ("let me talk to a manager")
- Out of scope: an irrelevant request, spam
How classification works: three approaches
Approach 1: Rules (keywords)
The simplest. A dictionary of keywords for each intent. "return", "refund" → return. "price", "how much" → a pricing FAQ. Fast, cheap, predictable. The downside: it's fragile. "I'd like to return to your services" will wrongly trigger a return.
Approach 2: AI classification (Claude Haiku)
You pass the message to the model and ask it to classify it. It's more accurate and understands context and sarcasm. It costs more than rules (you pay for tokens) and is a bit slower. In return, it makes noticeably fewer mistakes with real-world language. The cost per message is small and is calculated in tokens: current prices are on the What's current page.
Approach 3: Hybrid (recommended)
First check the rules: if one matches with high confidence, use the rule (cheap). If not, send it to Claude Haiku (accurate). This is the best balance of speed and cost for high-traffic bots.
Entity extraction: what's hidden in the text
Besides the intent, you need to pull out the specific details. A customer writes: "I want to order 3 boxes by Friday, delivered to 15 Main St." Here:
- Quantity: 3
- Deadline: Friday
- Address: 15 Main St
Without entity extraction, the bot answers "okay, we'll set that up" and loses all that information. With it, the bot pulls the details into the CRM, checks stock, and says whether you can make Friday.
For entity extraction we also use Claude or regex. Claude handles fuzzy phrasing better ("the day after tomorrow," "around five grand").
Escalation triggers: when to hand off to a human
This is the most important part of the system. Escalate too rarely and the customer leaves angry. Too often and your agent drowns in simple questions.
Smart escalation works off several signals:
1. An explicit request for an agent "I want to talk to a manager," "get me a human," "call me": the bot escalates immediately, no questions asked.
2. A frustration detector Analyze the tone: all caps, exclamation points, marker words ("nightmare," "terrible," "never again," "give me my money back"). Claude Haiku returns a sentiment score. If it's above the threshold, escalate.
3. A complex request The intent wasn't recognized twice in a row. The question needs information from several systems. A legal or financial question (per company policy).
4. A high-value lead The amount in the request is > $N. The customer asks about a business plan. They mentioned a competitor, which calls for a personal touch.
5. A system fallback After 3 failed attempts to answer, escalate automatically. Better to hand off to a human than to annoy the customer with an endless "I didn't understand."
State management: the conversation's memory
To store state we use Redis (fast, for active sessions) or Supabase (persistent, for history). Customer conversations contain personal data: store only what you need, limit how long you keep it, and don't send the model anything extra. In the state we store:
{
"session_id": "uuid",
"user_id": "telegram_chat_id",
"history": [...], # the last N messages
"current_intent": "complaint",
"entities": {"order_id": "12345"},
"frustration_score": 0.3,
"escalated": False,
"turns_without_resolution": 1
}With every new message we update the state and make the routing decision based on the whole history, not just the last line.
Multi-channel: one engine, many channels
The intent routing logic shouldn't be tied to Telegram or WhatsApp. The right architecture:
[Telegram] → adapter → [Intent Engine] → action → [Telegram response]
[WhatsApp] → adapter → [Intent Engine] → action → [WhatsApp response]
[Web Widget] → adapter → [Intent Engine] → action → [Web response]The adapter normalizes the incoming message into one format. The Intent Engine works the same way for everyone. The response is formatted for the right channel. This lets you support 3 channels with one codebase.
Alternatives: Botpress, Dialogflow, Rasa
Botpress: an open-source platform with a visual conversation builder. Good for complex conversation trees, with built-in NLU. Choose it if your team isn't technical and you need a visual editor.
Dialogflow (Google): an enterprise solution with powerful NLU. Integrates well with Google Workspace. More expensive, with vendor lock-in. Choose it for large corporate projects.
Rasa: open source with maximum flexibility. Requires ML expertise. Choose it if you need on-premise hosting and full control over your data.
Claude directly (our approach): the best choice for solo founders and small teams. Quick to start, low cost, flexible. Haiku for classification, Sonnet for complex answers.
Monitoring: three key metrics
Escalation rate: the percentage of conversations handed to an agent. There's no single standard: it depends on your niche and how complex the questions are. If almost every conversation goes to an agent, the bot is useless. If almost none do, the bot may not be escalating when it should.
Resolution rate: the percentage of questions the bot closes without an agent. Set your target based on your own pilot, not on other people's numbers. It grows as the knowledge base improves.
CSAT (Customer Satisfaction Score): the customer's rating after the conversation ends (1-5). Track it separately for bot sessions and agent sessions. The gap shows you where the weak spot is.
Build the monitoring dashboard in Grafana, or use a simple /stats command in a chat for the owner.
Code: an intent routing pipeline in Python + FastAPI
import os
import json
import redis
from fastapi import FastAPI, Request
from anthropic import Anthropic
app = FastAPI()
client = Anthropic()
r = redis.Redis(host='localhost', port=6379, decode_responses=True)
ESCALATION_PHRASES = [
"talk to a manager", "agent", "real person",
"human", "call me", "speak to someone"
]
FRUSTRATION_WORDS = [
"terrible", "nightmare", "never again", "scam",
"money back", "ripoff", "unacceptable"
]
def get_session(session_id: str) -> dict:
data = r.get(f"session:{session_id}")
if data:
return json.loads(data)
return {
"history": [],
"frustration_score": 0.0,
"turns_without_resolution": 0,
"escalated": False,
"current_intent": None
}
def save_session(session_id: str, state: dict):
r.setex(f"session:{session_id}", 3600, json.dumps(state))
def check_escalation_rules(message: str, state: dict) -> tuple[bool, str]:
"""Escalation rules: a quick check without AI."""
msg_lower = message.lower()
# An explicit request for an agent
if any(phrase in msg_lower for phrase in ESCALATION_PHRASES):
return True, "explicit_request"
# Frustration detector based on words
frustration_hit = sum(1 for w in FRUSTRATION_WORDS if w in msg_lower)
if frustration_hit >= 2:
return True, "frustration_detected"
# Frustration accumulated over the session
if state["frustration_score"] > 0.7:
return True, "cumulative_frustration"
# The bot failed to help 3 times in a row
if state["turns_without_resolution"] >= 3:
return True, "repeated_fallback"
return False, ""
def classify_intent(message: str, history: list) -> dict:
"""Claude Haiku classifies the intent and extracts entities."""
history_text = "\n".join([
f"{m['role']}: {m['content']}" for m in history[-4:]
])
response = client.messages.create(
model="claude-haiku-4-5", # check that the model is still available in the API; current models: the "What's current" page
max_tokens=300,
system="""You are an intent classifier for an online store's chatbot.
Return JSON with these fields:
- intent: one of [faq_price, faq_delivery, faq_return, complaint, purchase_intent, order_status, out_of_scope]
- confidence: 0.0-1.0
- entities: an object with the extracted data (order_id, amount, date, product)
- sentiment: positive/neutral/negative
- frustration_score: 0.0-1.0
JSON only, no explanations.""",
messages=[{
"role": "user",
"content": f"Conversation history:\n{history_text}\n\nNew message: {message}"
}]
)
try:
return json.loads("".join(b.text for b in response.content if b.type == "text"))
except Exception:
return {
"intent": "out_of_scope",
"confidence": 0.0,
"entities": {},
"sentiment": "neutral",
"frustration_score": 0.3
}
def generate_bot_response(intent: str, message: str, entities: dict, history: list) -> str:
"""Generate a response based on the intent."""
intent_prompts = {
"faq_price": "Answer the question about the product's price. If no specific product is mentioned, ask which one.",
"faq_delivery": "Answer about shipping: 2-5 days, free on orders over $50.",
"faq_return": "Answer about returns: 14 days, no questions asked, receipt required.",
"complaint": "Acknowledge the complaint with empathy. Ask for the details needed to resolve it.",
"purchase_intent": "Help with the purchase. Ask for details and offer to add it to the cart.",
"order_status": "Ask for the order number if it wasn't given.",
}
prompt = intent_prompts.get(intent, "Politely say you didn't understand the question and offer some options.")
response = client.messages.create(
model="claude-haiku-4-5",
max_tokens=200,
system=f"You are a polite assistant for an online store. {prompt}. Keep it short: 1-3 sentences.",
messages=[{"role": "user", "content": message}]
)
return "".join(b.text for b in response.content if b.type == "text")
async def notify_operator(session_id: str, message: str, reason: str, state: dict):
"""Notify the agent in Telegram."""
import httpx
bot_token = os.environ.get("TELEGRAM_BOT_TOKEN")
operator_chat_id = os.environ.get("OPERATOR_CHAT_ID")
history_preview = "\n".join([
f"{'👤' if m['role']=='user' else '🤖'} {m['content']}"
for m in state["history"][-3:]
])
text = (
f"🚨 Escalation | {reason}\n"
f"Session: {session_id}\n\n"
f"Recent messages:\n{history_preview}\n\n"
f"Latest: {message}"
)
async with httpx.AsyncClient() as http:
await http.post(
f"https://api.telegram.org/bot{bot_token}/sendMessage",
json={"chat_id": operator_chat_id, "text": text}
)
@app.post("/chat")
async def chat_endpoint(request: Request):
body = await request.json()
session_id = body["session_id"]
message = body["message"]
state = get_session(session_id)
# Quick check of the escalation rules
should_escalate, escalation_reason = check_escalation_rules(message, state)
if should_escalate and not state["escalated"]:
state["escalated"] = True
state["history"].append({"role": "user", "content": message})
save_session(session_id, state)
await notify_operator(session_id, message, escalation_reason, state)
return {
"response": "Got it. I'm connecting you with an agent, who will reply within 5 minutes.",
"escalated": True,
"reason": escalation_reason
}
# AI intent classification
classification = classify_intent(message, state["history"])
intent = classification.get("intent", "out_of_scope")
confidence = classification.get("confidence", 0.0)
# Update the state
state["current_intent"] = intent
state["frustration_score"] = max(
state["frustration_score"],
classification.get("frustration_score", 0.0)
)
# Low confidence: count it as unresolved
if confidence < 0.5 or intent == "out_of_scope":
state["turns_without_resolution"] += 1
else:
state["turns_without_resolution"] = 0
# Generate the response
bot_response = generate_bot_response(
intent, message,
classification.get("entities", {}),
state["history"]
)
# Save the history
state["history"].append({"role": "user", "content": message})
state["history"].append({"role": "assistant", "content": bot_response})
state["history"] = state["history"][-20:] # Keep the last 20 messages
save_session(session_id, state)
return {
"response": bot_response,
"intent": intent,
"confidence": confidence,
"escalated": False
}This code runs as a FastAPI service. A Telegram bot, a WhatsApp adapter or a website widget sends POST requests to /chat and gets back a response with the intent, confidence and an escalation flag.
Practice
Step 1: Set up the environment (5 min)
mkdir chatbot-manager && cd chatbot-manager
python -m venv venv && source venv/bin/activate
pip install fastapi uvicorn anthropic redis httpx python-dotenv
# .env file
echo "ANTHROPIC_API_KEY=your_key" >> .env
echo "TELEGRAM_BOT_TOKEN=your_bot_token" >> .env
echo "OPERATOR_CHAT_ID=your_chat_id" >> .env
# Start Redis locally (or use Redis Cloud)
docker run -d -p 6379:6379 redis:alpineStep 2: Set up a knowledge base for FAQs (10 min)
Create a knowledge_base.json file with answers to your business's common questions. The structure:
{
"faq_price": "Our products start at $10. Current catalog: shop.example.com/catalog",
"faq_delivery": "Shipping takes 2-5 business days. Free on orders over $50.",
"faq_return": "Returns within 14 days. Receipt and original packaging required. Refunds are issued within 3-5 days."
}Replace intent_prompts in the code with real answers from your knowledge base.
Step 3: Run it and test the routing (10 min)
uvicorn main:app --reload --port 8000Test with curl or Postman:
# A regular question
curl -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"session_id": "test1", "message": "how much is shipping?"}'
# An escalation test
curl -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"session_id": "test2", "message": "this is TERRIBLE, a total ripoff, I want my money back NOW!!!"}'
# An explicit request for an agent
curl -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"session_id": "test3", "message": "I want to talk to a real person"}'Check that the Telegram notification arrives when there's an escalation.
Step 4: Connect it to a Telegram bot (8 min)
# telegram_adapter.py
from telegram.ext import Application, MessageHandler, filters
import httpx, os
async def handle_message(update, context):
session_id = str(update.effective_chat.id)
message = update.message.text
async with httpx.AsyncClient() as client:
resp = await client.post(
"http://localhost:8000/chat",
json={"session_id": session_id, "message": message}
)
data = resp.json()
await update.message.reply_text(data["response"])
app = Application.builder().token(os.environ["TELEGRAM_BOT_TOKEN"]).build()
app.add_handler(MessageHandler(filters.TEXT, handle_message))
app.run_polling()The same /chat endpoint works for other channels: a WhatsApp or website widget adapter just sends the message in the same format.
Step 5: Set up monitoring (5 min)
Add a /stats endpoint for a quick look at the metrics in Redis:
@app.get("/stats")
async def get_stats():
keys = r.keys("session:*")
total = len(keys)
escalated = sum(1 for k in keys if json.loads(r.get(k)).get("escalated"))
return {
"total_sessions": total,
"escalated": escalated,
"escalation_rate": f"{escalated/total*100:.1f}%" if total else "0%",
"active_last_hour": total # simplified
}Tools and resources
| Tool | What it's for | Price |
|---|---|---|
| Claude Haiku | Intent classification, entity extraction | Pay per token; as of October 2026: $1 / $5 per million tokens (input / output) |
| Redis | Storing session state (fast) | Free self-hosted; cloud at the provider's rates |
| Supabase | Persistent conversation history | Has a free tier with limits |
| FastAPI | Backend for webhook endpoints | Free, open source |
| Botpress | Visual conversation builder (an alternative) | Has a free plan; pricing on its website |
| Dialogflow | Enterprise NLU (when you need a platform) | Pay as you go; pricing on the Google Cloud website |
| Rasa | On-premise, maximum control | Free, open source |
| Telegram Bot API | A channel + agent alerts | Free |
| Grafana | A metrics monitoring dashboard | Free self-hosted |
The recommended starter stack: FastAPI + Claude Haiku + Redis + Telegram. It comes together quickly and cheaply. Estimate the model cost like this: two Haiku calls per message (classification and response) × tokens × price per million tokens. Run a pilot on a hundred messages and look at usage in the API responses.
Key takeaways
"A bot without intent routing is a cashier who answers everything with 'combo number three.' A bot with routing is a dispatcher who knows where to send every request."
"The escalation rule is simple: when in doubt, hand it to a human. Better to spend an agent's time than to lose a customer over a bad bot answer."
"Escalation rate is your bot's honest KPI. If it's through the roof, the bot isn't working. If it's close to zero, the bot may be escalating too rarely and customers are leaving without a word."
Next lesson
→ Product analytics: PostHog, Mixpanel and smart insights
The mark stays in this browser only and is never sent anywhere. My progress