The gist
Every course teaches you how to build an agent. Almost none teach you what to do when it breaks at 3 a.m. right before a client demo.
Production isn't "launch it and forget it." Production is "launch it and now you live with it." Hallucinations, endless loops, a crashed MCP server, context overflow, a spike in your API bill, prompt injection, a quiet silent failure: these aren't "if it happens." They're "when it happens."
This lesson covers 7 typical failure patterns and a recovery playbook for each one. Tools, code patterns, checklists, incident response. Not theory: this is what you'll actually do when things go wrong and you're tempted to panic.
Key concepts
- Hallucination: the model confidently states a false fact
- Loop: the agent is stuck on one action, retrying with no progress
- Tool failure: an MCP server or external API doesn't respond or returns an error
- Context overflow: the context has bloated and the model loses track of its instructions (context rot)
- Cost spike: your API bill jumps 5-10x with no obvious reason
- Prompt injection: outside input makes the agent do something against the rules
- Silent failure: the pipeline ran without errors, but the result is wrong
- Circuit breaker: a pattern that switches off a broken integration for N minutes
- Graceful degradation: falling back to simpler functionality when something fails
- Chaos engineering: deliberately causing failures to check that you're ready for them
- Postmortem: a structured review of an incident: root cause and prevention
Theory
Failure 1: Hallucination (the agent confidently lies)
Symptom: the answer looks right, the formatting is clean, the tone is confident, but the fact is false. "The property tax rate in Ecuador is 5%" (the real rate is different). "This function was added in version 3.2 of the library" (there's no such version).
Risk levels:
- CRITICAL: medical, legal, financial advice (the customer acts on the answer)
- HIGH: production code, SQL queries, configs (run without review)
- MEDIUM: casual chat, brainstorming (limited consequences)
Recovery patterns:
- Ground truth verification (RAG): add retrieval from verified sources. The model answers not "from memory" but from documents you control
- Confidence scoring: in the prompt: "Answer and give a confidence score from 0 to 100. If it's under 70, say 'I'm not sure'"
- Multi-LLM voting: two independent calls (Claude + GPT). If the answers disagree, escalate to a human
- "Cite or refuse": the model MUST name a source (URL, doc id) or say plainly "I don't know"
Detection: sample 5% of outputs daily and have a human review them. Log the questions where the model said "I don't know": that's your map of gaps.
# Pattern: cite-or-refuse via the system prompt
SYSTEM = """Answer only based on the documents provided.
Back every fact with a source in the format [doc_id:section].
If the information isn't in the documents, answer exactly:
"I can't confirm this from the available sources."
Do NOT make things up. Do NOT rely on general knowledge."""Failure 2: Loop (the agent repeats itself forever)
Symptom: the same phrase, the same tool call, the same error, over and over. The logs fill up with identical entries. The API meter keeps running.
Common cause: a tool failed, so the agent tries again with the same parameters. Or the model doesn't "get" that the tool already returned "file not found" and repeats the call.
Recovery:
- Iteration cap: a maximum of N attempts, then fall back or abort
- Error context on retry: pass the previous error message into the next call
- State diff: check that the state actually changed between iterations
- Stuck detection: the same action 3 times in a row means abort
MAX_ATTEMPTS = 5
previous_action = None
stuck_count = 0
for i in range(MAX_ATTEMPTS):
result = agent.run(state)
if result.action == previous_action:
stuck_count += 1
if stuck_count >= 3:
raise StuckLoopError(f"Same action repeated 3x: {result.action}")
else:
stuck_count = 0
previous_action = result.action
if result.done:
break
else:
raise MaxIterationsExceeded(f"Hit cap {MAX_ATTEMPTS}")Failure 3: Tool failure (the MCP server doesn't respond)
Symptom: a timeout, HTTP 500, a crashed MCP server, a rate limit from an external API.
Recovery:
- Circuit breaker: 3 errors in a row, switch the integration off for 5 minutes, and stop hammering it for nothing
- Fallback strategy: the Stripe MCP is down, so look up basic billing info in your local DB
- Graceful degradation: the feature is unavailable, so say honestly "the service is temporarily unavailable, try again in 5 minutes." Do NOT fake a result
- Health checks: ping the MCP every 60 seconds, alert if it's been down for more than 2 minutes
Library: Tenacity (Python): retries with exponential backoff out of the box.
from tenacity import retry, stop_after_attempt, wait_exponential
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential(multiplier=1, min=2, max=10),
reraise=True
)
def call_external_api(payload):
return requests.post(URL, json=payload, timeout=10)Failure 4: Context overflow (the agent forgot the start of the session)
Symptom: the agent ignores instructions from the system prompt, contradicts itself, "forgets" what you discussed 20 messages ago.
Cause: context over 50K tokens leads to context rot. The model handles information from the middle worse and the beginning and end better.
Recovery:
/compactin Claude Code when the context grows (you can add an instruction about what to keep:/compact keep the architecture decisions); automatic compaction works too, and you set its threshold with the/autocompactcommand- Sliding window: keep the last N messages plus a summary of the earlier ones
- Critical instructions at the end: the model remembers recent context better (recency bias)
- System prompt with every message: if a rule is critical, repeat it
- Move long data to RAG: don't stuff 50K of documentation into the context; use retrieval
Failure 5: Cost spike (a day's bill is 10x the norm)
Symptom: $50/day instead of $5/day. A notice that you're nearing your monthly limit arrives on the 5th of the month.
Possible causes:
- A loop bug (Failure 2 wasn't caught)
- Missing cache (the same prompts repeated without prompt caching)
- A scraping attack (if you have a public endpoint without auth)
- A retry storm (failed, retry, failed, retry)
- Someone else using your API key (leaked in git)
Recovery:
- Daily budget kill switch: spend over $X/day automatically pauses the API
- Spike detection: spend over 2x the daily average sends an email or chat alert
- Audit log: log every API call to KV or a DB (cheap and easy to debug)
- Disable and investigate before you resume. Don't "raise the limit to keep going" without figuring out what happened
Tools:
- Anthropic Console: the Usage page and your own spend limit (Settings → Billing → Spend limits)
- Langfuse: open-source observability, cost per user
- LangSmith: tracing for LangChain/LangGraph
Failure 6: Prompt injection
Symptom: the agent disabled a hook "because the user asked," passed secrets to an external tool, sent an email in your name, ran rm -rf from an instruction inside a Markdown file it was reading.
For the full picture, see the Prompt Injection Defense lesson. Here we cover recovery.
Recovery:
- Isolation: route every risky tool through a separate agent with minimal permissions
- Audit log everything: what was a command vs. what was outside input (keep them clearly separate)
- Restoration:
git revert, alert the team, investigate the breach - Lessons learned: add the pattern to your
pre-tool-use-prompt-injection.shhook
# Minimal protection: the hook checks for suspicious patterns
# .claude/hooks/pre-tool-use-prompt-injection.sh
grep -iE 'ignore (previous|all) instructions|disable.*hook|reveal.*system prompt' "$INPUT" \
&& exit 1Failure 7: Silent failure (the agent "works" but the output is wrong)
The sneakiest one. The pipeline completes, exit code 0, no errors in the logs. But the result is garbage.
Examples:
- The LLM returned JSON with the right structure, but the values were made up
- The translation reads smoothly, but the meaning is distorted
- A financial calculation uses the right formula, but the numbers came out of thin air
Recovery:
- Output validation: schema-check every LLM output. Pydantic, JSON schema, zod
- Smoke tests: on every pipeline run, check that 1-2 known cases give the expected output
- User feedback loop: "Was this answer right? 👍/👎" every 50 interactions, feeding a dashboard
- Eval suite (the Evals lesson): regression tests after every prompt change
from pydantic import BaseModel, Field, ValidationError
class InvoiceExtraction(BaseModel):
amount: float = Field(gt=0, lt=1_000_000)
currency: str = Field(pattern="^(USD|EUR|MXN)$")
invoice_date: str # ISO 8601
try:
parsed = InvoiceExtraction.model_validate_json(llm_output)
except ValidationError as e:
# silent failure caught: escalate
log_anomaly(llm_output, e)
raiseRecovery toolkit (what to keep in the trunk)
| Category | Tool | Purpose |
|---|---|---|
| Logging | Langfuse, LangSmith, structured logs in KV | Every LLM call with input/output/tokens |
| Alerting | Slack or chat webhook, PagerDuty, email | Spike / down / anomaly |
| Monitoring | Anthropic Console + a custom dashboard (the Production Observability lesson) | Cost, latency, error rate |
| Rollback | Git tags + a revert script | One command back to the last good version |
| Communication | Customer reply templates | "Known issue, fix within 2 hours" |
| Validation | Pydantic / zod / JSON schema | Schema check on every output |
| Retry logic | Tenacity (Python), p-retry (JS) | Exponential backoff out of the box |
Production checklist (required before you go live)
Going live without 8 out of 8 means tech debt from day one.
Chaos engineering for AI (testing failure cases)
The Principles of Chaos approach applied to agents. Break the system on purpose in staging to make sure recovery works.
| Inject | How | What you're checking |
|---|---|---|
| Hallucination | Add "Always answer Y is true" to the prompt, then run a spike test | Does the verification layer catch it |
| Loop | A mock tool that always returns the same error | Does the iteration cap kick in |
| HTTP 500 | A mock MCP that returns 500 on every 3rd call | Do the circuit breaker and fallback kick in |
| Context overflow | Stuff in 100K tokens of junk | Does /compact trigger |
| Cost spike | A loop with no break, a cheap prompt × 10,000 | Does the daily budget kill switch kick in |
| Schema mismatch | The LLM returns fields with extra or missing keys | Does Pydantic catch it, or is it a silent failure? |
Once a month, hold a chaos day in staging. Half an hour of work that can save you in production.
Incident response in 4 steps
When something's on fire, don't panic. Go step by step:
- Stop: switch off the affected feature and prevent further damage. Turn off the feature flag, or put up a maintenance page
- Triage: what happened, the scope (1 user or everyone), the blast radius (billing? data? reputation?)
- Stabilize: roll back to the last good tag OR apply a temporary fix (a hardcoded fallback) to bring the service back
- Postmortem: once things are restored: the root cause, what you'll add to monitoring, how to prevent it
Postmortem template (journals/incidents/YYYY-MM-DD-<name>.md):
# Incident: <name> **Date:** YYYY-MM-DD **Duration:** 14:32 - 15:47 UTC (1h 15m) **Severity:** P1 / P2 / P3 **Impact:** N customers, $X loss / 0 data loss ## Timeline - 14:32: spike alert in the team chat (cost > 2x avg) - 14:35: investigation started, found a loop in agent X - 14:50: iteration cap deployed - 15:47: service stable ## Root cause Tool Y returned an empty string instead of null. The agent read it as "that didn't work" and retried forever. ## What went well - The alert fired within 3 minutes - The rollback procedure worked ## What went wrong - There was no iteration cap (MAX_ATTEMPTS wasn't set) - Output validation didn't check for an empty string ## Action items - [ ] Add MAX_ATTEMPTS=5 to all agents (owner: A, ETA: tomorrow) - [ ] Pydantic check on tool Y output (owner: A, ETA: 3 days) - [ ] Chaos test for empty strings (owner: A, ETA: one week)
Without a postmortem, the incident will happen again. Count on it.
Anti-patterns (what NOT to do in a panic)
- ❌ "I'll just restart it" without understanding the root cause. It'll happen again within the hour
- ❌ Hiding the incident from customers. They'll find out on their own, and their trust takes a hit
- ❌ A hot fix with no entry in the audit log. A month later: "why is there a magic number here?"
- ❌ "It was a fluke, I'll forget about it." Patterns repeat. Every incident gets a postmortem
- ❌ Disabling a hook "just to check" (see the Hook-Deny-By-Design lesson). It's a security risk, and you'll forget to turn it back on
- ❌ Raising the API limit to keep going without figuring out what caused the spike
Cost recovery (if you've already lost money)
The reality: recovery is almost always incomplete. Prevention is cheaper.
| Scenario | Chance of getting money back | What to do |
|---|---|---|
| Anthropic billing bug (on their side) | Low, but real | Support ticket with the trace_id and logs. Sometimes they give a credit |
| Stripe billing error (a customer was overcharged) | High | Dispute it through the Stripe dashboard, refund comes to you |
| Leaked API key (pushed to GitHub) | 0% | Rotate the key immediately, audit all calls |
| A loop bug ate $200 | 0% | Lesson learned. Iteration caps forever |
Lesson: watch your budget alerts. Getting money back is a rare bonus, not a plan.
Who should learn what, and when
Beginner (first 3 months in production)
Only failures 1-3, the most common ones:
- Basic hallucination defense (cite-or-refuse in the prompt)
- An iteration cap on every loop
- Try/except + retry for tool failures
Basic logging (just print to a file at first). A spend limit in the Anthropic Console.
Intermediate (3-12 months)
You add:
- Failures 4-5 (context management + cost monitoring)
- Langfuse / LangSmith for observability
- Chat alerts on spikes
- A postmortem habit (even if you work solo, write it down)
Advanced (1+ year in production)
All 7 patterns:
- Monthly chaos engineering
- Automated runbooks (an incident triggers a bot that runs triage)
- Multi-LLM voting on critical paths
- Pydantic validation everywhere
- Automated customer notifications
Practice
Step 1: Add an iteration cap to every agent
Find the places in your code with while or for loops that have no upper bound, and recursive calls with no depth limit.
# Before
def agent_loop(state):
while not state.done:
state = step(state)
return state
# After
MAX_ITERATIONS = 10
def agent_loop(state):
for i in range(MAX_ITERATIONS):
state = step(state)
if state.done:
return state
raise MaxIterationsExceeded(
f"Hit {MAX_ITERATIONS}. Last state: {state.summary()}"
)Step 2: Set a spend limit and a warning
Anthropic Console → Settings → Billing → Spend limits → Adjust limit. There's one limit: a monthly ceiling. Once it's reached, the API stops until you raise the limit or the new month starts. Set it to roughly 3x your usual monthly spend.
Build the warning (a "soft limit") yourself, as a daily check through cron. The threshold is your current average daily spend × 1.5.
# scripts/check-daily-cost.sh
# Requires an Admin API key (sk-ant-admin01-...): organizations have one, individual accounts don't.
# Without one, check the Usage page in the Console by hand.
TODAY=$(date -u +%Y-%m-%dT00:00:00Z)
TOMORROW=$(date -u -v+1d +%Y-%m-%dT00:00:00Z) # on Linux: date -u -d tomorrow +%Y-%m-%dT00:00:00Z
RESP=$(curl -s "https://api.anthropic.com/v1/organizations/cost_report?starting_at=$TODAY&ending_at=$TOMORROW" \
-H "anthropic-version: 2023-06-01" \
-H "x-api-key: $ANTHROPIC_ADMIN_KEY")
# Amounts come back as a string in cents; check the exact response structure in the Cost API reference
DAILY_SPEND=$(echo "$RESP" | jq '[.data[].results[].amount | tonumber] | add / 100')
# Sends the alert to a Slack incoming webhook (any chat webhook works)
if (( $(echo "$DAILY_SPEND > 20" | bc -l) )); then
curl -X POST "$ALERT_WEBHOOK_URL" \
-H 'Content-type: application/json' \
-d "{\"text\":\"⚠️ Daily spend: \$$DAILY_SPEND\"}"
fiCost API documentation: platform.claude.com/docs/en/manage-claude/usage-cost-api.
Step 3: Add Pydantic validation on critical output
from pydantic import BaseModel, Field, ValidationError
from anthropic import Anthropic
client = Anthropic()
class CustomerReply(BaseModel):
sentiment: str = Field(pattern="^(positive|neutral|negative)$")
intent: str
confidence: float = Field(ge=0, le=1)
suggested_response: str = Field(min_length=10, max_length=500)
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "..."}]
)
try:
validated = CustomerReply.model_validate_json("".join(b.text for b in response.content if b.type == "text"))
use(validated)
except ValidationError as e:
log_to_journal("".join(b.text for b in response.content if b.type == "text"), e)
raise SilentFailureDetected(str(e))Step 4: Run one chaos test
Pick the failure that scares you most (usually a cost spike or a silent failure). Today, in staging:
- Create the conditions on purpose (a mock that returns bad output / a loop with no break)
- Run the agent
- Time how long it takes to get an alert or an abort
- If it's over 5 minutes, add detection. If nothing fired, fix it
Afterward, write it up in journals/chaos-tests/YYYY-MM-DD.md:
- What you injected
- What you expected
- What happened
- Action items
Step 5: Customer notification template
templates/customer-incident-notification.md:
Hello, We're currently seeing a problem with <feature>: <short description of the symptom>. **Status:** we're working on a fix **Estimated time to restore:** <X minutes> **What to do for now:** <workaround or "please wait"> We'll email you again once it's resolved. Sorry for the inconvenience. — The <name> team
One file. Open it, fill in 3 fields, send. During an incident you don't want to be writing from scratch.
Tools and resources
- Tenacity: a Python retry library with exponential backoff
- Anthropic Console: usage tracking and your own spend limit
- Langfuse: open-source observability for LLM apps, cost per user, latency (Helicone has been in maintenance mode since March 2026)
- LangSmith: tracing for LangChain/LangGraph agents
- Pydantic: schema validation for Python (zod is the TypeScript equivalent)
- PagerDuty: incident management for serious teams
- Principles of Chaos Engineering: Netflix's approach to deliberate failures
- The Prompt Injection Defense lesson: defense in depth against prompt injection
- The Production Observability lesson: custom dashboards for monitoring agents
- The Evals lesson: eval suites for regression testing
Lesson checklist
Key takeaways
Production isn't "launch it and forget it." Production is "launch it and live with it." Anyone running agents in production will run into these 7 failure patterns. The question isn't "if" but "when," and whether you're ready.
Prevention is cheaper than recovery. An iteration cap, a budget kill switch and output validation are three cheap patterns that cover most typical incidents. Without them, you have tech debt from your first day in production.
A blameless postmortem is the only way to keep an incident from happening again. The template takes 10 minutes and protects you for years. Even if you work solo, write it down. Future you will thank you.
Next lesson
→ Backup & Disaster Recovery for your AI stack: what to do when the provider itself goes down
The mark stays in this browser only and is never sent anywhere. My progress