The gist
Picture a smart front desk at a big hospital. Every patient goes there first. The front desk handles most questions on its own: how to book an appointment, where to park, what paperwork to bring. For some cases it writes a short summary and hands the patient to the right doctor. Urgent cases get flagged to the doctor on call right away. Doctors spend their time only on what actually needs a doctor.
AI Customer Support works the same way. It doesn't replace the support team; it works alongside it. It takes the routine questions and frees people up for the hard ones.
In this lesson we build a complete system: a ticket system with classification, RAG over the company's knowledge base, automatic replies and smart escalation. The result is a foundation you can turn into a service for clients.
Key concepts
- RAG (Retrieval-Augmented Generation): Claude answers based on specific documents instead of making things up
- Ticket classification: routine / unusual / urgent
- Escalation: automatically handing hard cases to a person
- Deflection Rate: the percentage of questions AI resolved without a person
- CSAT: the customer's rating (Customer Satisfaction Score)
Theory
Why AI Customer Support
Without AI, the first reply from support often takes several hours. That's fine for a hard question. But most questions are routine: "how do I change my plan," "where's my package," "how do I connect an integration." For those, the customer waits for hours and gets a two-line answer the agent copied from the FAQ.
With AI:
- Routine question → automatic reply in 30 seconds (happy customer)
- Unusual question → AI draft for the agent, the agent edits it → 5 minutes instead of 20
- Urgent question → immediate escalation + a notification to the agent
Now the agent only handles unusual and urgent cases. Their workload drops noticeably, and the quality of answers to hard questions goes up, because they can focus.
System architecture
Customer writes a question
↓
Classification (Claude): routine / unusual / urgent
↓
Routine → RAG over the knowledge base → automatic reply (most of them, roughly 80%)
Unusual → AI draft for the agent → agent edits → sends (roughly 15%)
Urgent → immediate escalation + notification → agent replies personally (roughly 5%)Three layers:
- Classifier: decides where to route the question
- RAG over the knowledge base: finds the exact answer in the documents
- Channel integration (Intercom / Telegram / email): receives the question and sends the reply
RAG: a cheat sheet the AI never forgets
RAG stands for Retrieval-Augmented Generation. It sounds complicated, but the idea is simple.
Regular Claude answers from its general knowledge. That works for general questions. It doesn't work when you need an answer about a specific company: its plans, its refund policy, the details of its product.
RAG means you give Claude a cheat sheet right in the prompt. FAQ documents, instructions, policies: all of it goes into the system prompt or the context. Claude sees the knowledge base and answers strictly from it. If the answer isn't in the base, it says so honestly instead of making something up.
Think of it this way: not a student trying to remember, but a student with a cheat sheet. The answer is accurate. It cites the source. If it's not on the cheat sheet, the student says so.
Practice
Step 1. Knowledge base structure
Create a knowledge-base/ folder with MD files. One file per topic:
knowledge-base/
billing.md — billing questions
shipping.md — shipping and returns
integrations.md — connecting integrations
plans.md — pricing plans
account.md — account managementExample billing.md:
## How to change your plan
Go to your account → Settings → Plan → choose a new plan.
The change takes effect immediately. Any extra charge is prorated.
## Refunds
Refunds are available within 14 days of payment on your first order.
To request a refund, email [email protected] with the subject "Refund" and your order number.
Processing takes 5 business days.
## Differences between the Basic and Pro plans
Basic: up to 5 users, 10 GB of storage, email support.
Pro: unlimited users, 100 GB of storage, priority support, API access.Step 2. Loading the knowledge base and answering a ticket
import anthropic
import json
import time
from pathlib import Path
from datetime import datetime, timezone
client = anthropic.Anthropic()
def load_knowledge_base(kb_path: str) -> str:
"""Loads the knowledge base from MD files"""
kb_content = []
for md_file in Path(kb_path).glob("**/*.md"):
kb_content.append(f"\n## {md_file.stem}\n{md_file.read_text()}")
return "\n".join(kb_content)
KB = load_knowledge_base("knowledge-base/")
def answer_support_ticket(question: str) -> dict:
"""Classifies the ticket and either answers or escalates"""
start_time = time.time()
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=800,
system=f"""You are an AI customer support agent.
COMPANY KNOWLEDGE BASE:
{KB}
Rules:
1. If the answer is in the knowledge base, answer accurately using it and don't add anything of your own
2. At the end of the answer, write: Source: [name of the knowledge base section]
3. If the answer is not in the knowledge base, write ESCALATE on the first line and explain why a person is needed
4. If the question is about a refund, ALWAYS write ESCALATE (even if the answer is in the base)
5. If the question is about a technical problem on the customer's side, ESCALATE
6. Tone: friendly, specific, no filler""",
messages=[{"role": "user", "content": question}]
)
answer = "".join(b.text for b in response.content if b.type == "text")
response_time = time.time() - start_time
return {
"answer": answer,
"needs_human": answer.strip().startswith("ESCALATE"),
"response_time_sec": round(response_time, 2),
"confidence": "low" if answer.strip().startswith("ESCALATE") else "high"
}
def track_metrics(result: dict, question: str):
"""Writes metrics to the log"""
log_entry = {
"ts": datetime.now(timezone.utc).isoformat(),
"escalated": result["needs_human"],
"response_time_sec": result["response_time_sec"],
"confidence": result["confidence"],
"question_length": len(question)
}
with open("support-metrics.jsonl", "a") as f:
f.write(json.dumps(log_entry, ensure_ascii=False) + "\n")
# Usage example
if __name__ == "__main__":
questions = [
"How do I change my plan?",
"I want a refund for my subscription",
"My Slack integration isn't working, everything is broken"
]
for question in questions:
print(f"\nQuestion: {question}")
result = answer_support_ticket(question)
track_metrics(result, question)
if result["needs_human"]:
print(f"ESCALATION -> handing off to a live agent")
print(f"Reason: {result['answer']}")
else:
print(f"Automatic reply ({result['response_time_sec']} sec):")
print(result["answer"])Step 3. Intercom integration
from flask import Flask, request
import requests
import os
app = Flask(__name__)
INTERCOM_TOKEN = os.environ["INTERCOM_TOKEN"]
AI_BOT_ID = os.environ["INTERCOM_BOT_ID"]
def assign_to_human_agent(conversation_id: str, priority: str = "normal"):
"""Assigns the ticket to a live agent and adds a tag"""
requests.post(
f"https://api.intercom.io/conversations/{conversation_id}/parts",
headers={
"Authorization": f"Bearer {INTERCOM_TOKEN}",
"Content-Type": "application/json"
},
json={
"type": "admin",
"admin_id": AI_BOT_ID,
"message_type": "assignment",
"assignee_id": None # assigns to the team, not a specific agent
}
)
# Add the priority tag
if priority == "high":
requests.post(
f"https://api.intercom.io/conversations/{conversation_id}/tags",
headers={"Authorization": f"Bearer {INTERCOM_TOKEN}"},
json={"id": os.environ["INTERCOM_HIGH_PRIORITY_TAG_ID"]}
)
@app.post("/intercom-webhook")
def handle_message():
data = request.json
if data.get("type") != "conversation.user.created":
return {"status": "ignored"}
item = data["data"]["item"]
conversation_id = item["id"]
parts = item["conversation_parts"]["conversation_parts"]
if not parts:
return {"status": "no_message"}
message = parts[0]["body"]
# AI answers
result = answer_support_ticket(message)
track_metrics(result, message)
if not result["needs_human"]:
# Send the automatic reply as the bot
requests.post(
f"https://api.intercom.io/conversations/{conversation_id}/reply",
headers={
"Authorization": f"Bearer {INTERCOM_TOKEN}",
"Content-Type": "application/json"
},
json={
"type": "admin",
"admin_id": AI_BOT_ID,
"message_type": "comment",
"body": result["answer"]
}
)
else:
# Let the customer know a person is joining
requests.post(
f"https://api.intercom.io/conversations/{conversation_id}/reply",
headers={
"Authorization": f"Bearer {INTERCOM_TOKEN}",
"Content-Type": "application/json"
},
json={
"type": "admin",
"admin_id": AI_BOT_ID,
"message_type": "comment",
"body": "Your question has been passed to a specialist. We'll reply within 2 hours."
}
)
assign_to_human_agent(conversation_id, priority="high")
return {"status": "ok"}
if __name__ == "__main__":
app.run(port=5000)Step 4. A support chat bot (Telegram example)
If the client doesn't use Intercom, a chat bot can cover the job in a couple of hours. The example below uses Telegram because its bot API is free and simple; the same pattern (receive a message, answer or escalate, notify the team) applies to other messaging channels your client's customers use.
import asyncio
import os
from telegram import Update, Bot
from telegram.ext import Application, MessageHandler, filters
TELEGRAM_BOT_TOKEN = os.environ["TELEGRAM_BOT_TOKEN"]
SUPPORT_TEAM_CHAT = os.environ["SUPPORT_TEAM_CHAT_ID"]
bot = Bot(token=TELEGRAM_BOT_TOKEN)
async def handle_support_message(update: Update, context):
user_message = update.message.text
user_id = update.effective_user.id
username = update.effective_user.username or str(user_id)
# Show that the bot is working
await update.message.reply_text("Checking...")
result = answer_support_ticket(user_message)
track_metrics(result, user_message)
if not result["needs_human"]:
await update.message.reply_text(result["answer"])
else:
# To the customer: a message that a person is joining
await update.message.reply_text(
"Your question needs a specialist's attention. "
"We'll reply within 2 business hours."
)
# To the team: a notification with the full context
escalation_text = (
f"Escalation from @{username} (id: {user_id})\n\n"
f"Question: {user_message}\n\n"
f"Escalation reason: {result['answer']}"
)
await bot.send_message(
chat_id=SUPPORT_TEAM_CHAT,
text=escalation_text
)
def run_bot():
application = Application.builder().token(TELEGRAM_BOT_TOKEN).build()
application.add_handler(
MessageHandler(filters.TEXT & ~filters.COMMAND, handle_support_message)
)
application.run_polling()
if __name__ == "__main__":
run_bot()Step 5. Metrics
Without metrics, you don't know whether the system works. Four numbers to track:
| Metric | What it measures | Target |
|---|---|---|
| First Response Time | Time to the first reply | < 1 min (AI), < 4 h (human) |
| Deflection Rate | % of tickets resolved without a person | > 75% |
| Resolution Rate | % of tickets closed with the first reply | > 60% |
| CSAT Score | Customer rating, 1–5 | > 4.2 |
def generate_support_report(metrics_file: str = "support-metrics.jsonl") -> dict:
"""Calculates the main metrics for the period"""
entries = []
with open(metrics_file) as f:
for line in f:
entries.append(json.loads(line))
if not entries:
return {"error": "no data"}
total = len(entries)
escalated = sum(1 for e in entries if e["escalated"])
deflection_rate = round((total - escalated) / total * 100, 1)
avg_response_time = round(
sum(e["response_time_sec"] for e in entries) / total, 2
)
return {
"total_tickets": total,
"escalated": escalated,
"auto_resolved": total - escalated,
"deflection_rate_pct": deflection_rate,
"avg_response_time_sec": avg_response_time
}
# Example output:
# {
# "total_tickets": 150,
# "escalated": 28,
# "auto_resolved": 122,
# "deflection_rate_pct": 81.3,
# "avg_response_time_sec": 2.4
# }Step 6. Updating the knowledge base
A knowledge base goes stale. New plans, new features, a changed refund policy. A simple update process:
- Edit the MD file in
knowledge-base/ - Restart the server (or add hot-reload)
- Claude immediately answers based on the new data
This is the main advantage of MD files over vector databases: an update takes 30 seconds, with no reindexing. The approach works as long as the knowledge base fits in the model's context. If the base is large and you get a lot of questions, turn on prompt caching (see the lesson Prompt Caching and the Batch API): the repeated part of the prompt costs less.
Tools and resources
| Tool | What it's for | Price |
|---|---|---|
| Intercom | Main ticket system (enterprise) | Paid, plans on its website |
| Crisp | An alternative to Intercom (simpler and cheaper) | Plans on its website |
| Telegram Bot API | Free support channel | Free |
| Flask | Webhook server for integrations | Free |
| Python-telegram-bot | Telegram bot library | Free |
| Claude Sonnet | Main model (balance of price and quality) | Depends on the size of the knowledge base and the reply; see What's current |
Minimal MVP stack:
- Python 3.11+
anthropic: the Claude SDKpython-telegram-bot: if you use Telegramflask: if you use a webhook integration- MD files: the knowledge base
Build-to-Sell
You can package this as a service for small companies where 1–3 people handle support. Whether you'll find clients and how much they'll pay depends on the niche, the market and your work. There are no guarantees.
The economics for the client:
Work it out together with the client, using their numbers. Hours per day spent on routine questions × the employee's hourly rate = what those questions cost per day. AI frees up only part of those hours: hard tickets and reviewing answers stay with people. Payback period = price of the service ÷ actual savings per day. An example with made-up numbers: 4 hours × $15 = $60 per day before AI, and if AI handles half the questions, the savings are about $30 per day.
How to structure the offer:
Set prices based on your costs and the value to the client; more in the lesson How to set a price.
| Option | What it includes |
|---|---|
| Setup | Development + configuration + the first knowledge base |
| Monthly support | Hosting + monitoring + knowledge base updates |
| Enterprise | Custom integration + team training |
MVP development time: 4–6 hours (chat bot + Claude + an MD FAQ).
What you sell the client:
- A chat bot or an Intercom/Crisp integration
- A knowledge base set up from their FAQ
- A metrics dashboard (Deflection Rate, response time)
- Documentation on updating the knowledge base
Key takeaways
- AI Customer Support doesn't replace the team. It strengthens it: AI takes the routine work, people keep the hard cases
- 80/15/5 is a guideline, not a law: most questions get an automatic reply, fewer get an AI draft for the agent, the rest are escalated; every company's proportions are different
- RAG with MD files is the simplest path to accurate answers. Updating the knowledge base takes 30 seconds
- Three metrics that matter to the client: Deflection Rate (> 75%), First Response Time (< 1 min), CSAT (> 4.2)
- A minimal MVP comes together in a few hours. As a service, you can package it in three parts: setup, monthly support and custom integrations
Next lesson
→ Call Support AI: Vapi + Bland.ai, voice calls in support
We look at the next level: not text tickets but voice calls. How AI picks up a call, answers from the knowledge base and hands hard cases to a person.
The mark stays in this browser only and is never sent anywhere. My progress