Library · Reliability: monitoring, failures, backups

Production observability: what to monitor when your agent is live

Engineer60 minUpdated: October 2026
89 of 105 in the library

Time: about 25 min reading + 35 min practice


The gist

"It works on my machine" is the most expensive phrase in the industry. You deployed an agent, it works, customers are using it. And then a message arrives: "it's been broken since last night." And you realize you found out not from your own system but from an annoyed user.

Production observability is your eyes and ears inside a running system. Without it you're a blind pilot: the engine is running, but you don't know when it will overheat, when the fuel will run out, or when a wing will fall off.

This lesson covers 5 metrics you MUST monitor, 3 levels of alerting and 4 tools for AI agents (the list as of October 2026). No fluff, with example numbers and dashboard mockups. The thresholds in this lesson are a starting point: over time, tune your own based on real data.

🎨 Picture this: a plane with no instrument panel. The engine is running and you're flying. But the engine temperature is rising. The fuel is running out. You're losing altitude. Without instruments, you'll only find out when something catches fire. Observability is your AI agent's instrument panel. It's not a luxury; it's what makes flying possible at all.


🎯 The main principle: signals, not logs

Logging ≠ observability. You can have gigabytes of logs and still not understand what's going on. Observability is the ability to answer questions about the system without digging into the code.

The three classic pillars (per the OpenTelemetry standard https://opentelemetry.io):

  1. Metrics: numbers over time (latency, cost, errors)
  2. Traces: the path of a single request through the whole system
  3. Logs: events with context

For AI agents, a fourth gets added:

  1. LLM-specific: prompts, completions, token usage, cache hits

🎨 Picture this: metrics are a patient's pulse and temperature (the overall condition). Traces are an X-ray of a specific bone (the detailed path of one problem). Logs are the chart with the patient's medical history. LLM-specific data is a blood test for markers specific to an AI patient.


Key concepts

  • p50 / p95 / p99: latency percentiles: half of requests are faster than p50, 95% faster than p95, 99% faster than p99. p99 shows the worst case your least well-served users see
  • Cardinality: the number of unique values in a metric. High-cardinality metrics (per user) are expensive to store
  • Sampling: recording not 100% of traces but 1-10% of normal ones + 100% of errors
  • Alert fatigue: having so many alerts that you ignore all of them. More dangerous than having no alerts
  • SLI / SLO / SLA: Service Level Indicator (the metric), Objective (the internal target), Agreement (the contract with the client)
  • Anomaly detection: a deviation from the baseline (yesterday's average × 2 = suspicious)
  • Trace ID: a unique request ID that passes through every component, for correlation
  • Cache hit ratio: the % of requests that hit the prompt cache. It directly affects cost

Theory

5 mandatory metrics

This is the minimum. Without them you don't understand what's happening with your agent in production.


Metric 1: Latency (response time)

What to measure: p50, p95, p99 response time per request

Targets for 2026:

Agent type p50 p95 p99
Chat (text) <3s <8s <15s
Voice (real-time) <500ms <1s <2s
Batch (background) OK 1-24h OK 1-24h OK 1-24h
API (B2B integration) <1s <3s <5s

Why p99 matters more than p50: if you have 1,000 requests a day and p99 = 30s, that means 10 users a day wait half a minute. They leave. A p50 of 2s looks nice in a report but hides the problem.

Tools for tracking it:

  • LangSmith (https://www.langchain.com/langsmith): tracing for LLM calls, with a free plan (limits on the website)
  • Helicone (https://www.helicone.ai/): an LLM proxy with metrics; in March 2026 the company was acquired by Mintlify, and the service is in maintenance mode with no new features
  • Custom Prometheus/Grafana: self-hosted, free but a time sink to set up

🎨 Picture this: p50 is the average time stuck in traffic for most drivers. p99 is how long the unluckiest one waits. If the average is 5 minutes but 1% of drivers sit for an hour, that's a broken traffic light nobody reports to you because "overall it's fine."


Metric 2: Cost per interaction

What to measure:

  • $/conversation (one user session)
  • $/user (cumulative for the month)
  • $/feature (where the money goes)
  • Total monthly burn (the overall bill)

Anomaly threshold: a spike >2x the daily average → alert. If you spent $20 yesterday and you're already at $80 by lunch today, something's wrong. Possibly:

  • A bug introduced an infinite loop of calls
  • Someone found a way to abuse it (prompt injection, repeated requests)
  • The code accidentally switched the model from Haiku to Opus

Tools:

  • Anthropic Console (https://platform.claude.com): the built-in billing dashboard
  • LangSmith / Helicone: the cost of each call, broken down by user and feature
  • A custom KV logger: a homemade counter in Cloudflare KV / Redis

Cost-saving signals that monitoring reveals:

  • A low cache hit ratio (<30%) → rethink your prompt structure
  • Output tokens growing faster than input → the agent has gotten chattier (rethink the system prompt)
  • Opus calls >20% of the total → check whether you're using Opus where Sonnet would do

Metric 3: Error rate

What to measure (4 categories):

  1. API errors: 4xx, 5xx from Anthropic/OpenAI
  2. Tool failures: MCP timeouts, schema mismatches, tool exceptions
  3. Validation failures: the LLM output doesn't match the expected JSON schema
  4. User-reported errors: explicit feedback saying "it doesn't work"

Target: <1% error rate in a steady state, <5% during a spike

Alert thresholds:

  • Error rate >5% over 5 minutes → page on-call (something broke right now)
  • Error rate >2% over 1 hour → warning (a degradation trend)
  • A single error code >50 events/hour → investigate (a systemic problem)

Categorizing matters: different errors call for different responses:

Error type Action
529 overloaded_error (Anthropic) Retry with backoff; not our fault
400 invalid_request_error A bug in the code; fix immediately
429 rate_limit_error Move up a tier or throttle
tool_use schema mismatch The LLM is unstable; add validation
JSON parse error Rework the system prompt

🎨 Picture this: the error rate is a patient's temperature. 98.6°F (1%) is normal. 100°F (2%) is a cold. 102°F (5%) means get to a doctor right away. And it's not just about seeing the number; you need to figure out whether it's the flu, an allergy or appendicitis.


Metric 4: Token usage

What to measure:

  • Input tokens / output tokens (per request + aggregate)
  • Cache hit ratio (% that hit the prompt cache)
  • Per-model breakdown (% Haiku / % Sonnet / % Opus)
  • Context length distribution (you see when people hit the limit)

Anomaly signals:

Signal What it means
Input tokens >5x average for one user A loop bug, scraping, a prompt injection attack
Cache hit <20% The prompt structure isn't optimized
Average output tokens growing for a week The system prompt is degrading; the agent has gotten chattier
Context >150K on a regular basis Time for a compaction strategy

Cache hit ratio is a critical metric: reading from the prompt cache costs about 10% of the regular input token price (as of October 2026; current terms: What's current). If the cache hit ratio is high, most of your input costs several times less. If the cache hit ratio = 0%, you pay the full rate.

python
# Tracking the cache hit ratio
def log_request(response):
    cache_tokens = response.usage.cache_read_input_tokens or 0
    total_input = response.usage.input_tokens + cache_tokens
    cache_ratio = cache_tokens / total_input if total_input > 0 else 0

    metrics.record("cache.hit_ratio", cache_ratio)
    metrics.record("tokens.input", response.usage.input_tokens)
    metrics.record("tokens.output", response.usage.output_tokens)

Metric 5: Business metrics

Technical metrics show that the system works. Business metrics show that the system works for the business.

What to measure:

Metric Description Target
Conversation completion rate % of sessions that reached their goal >70%
CSAT (customer satisfaction) Explicit feedback, 1-5 stars >4.0
Escalation rate % of "get me a human" cases <15%
Conversion rate For sales/marketing AI Varies
Time to resolution How many turns until it's resolved <5
Retention % of users who come back >30% in week 2

Tools:

  • PostHog (https://posthog.com): product analytics, with a free monthly allowance (see its pricing page for the current size)
  • Mixpanel (https://mixpanel.com): funnel analysis; pricing on the website

Why business metrics matter more than technical ones: you can have p95 = 1s and a 0.5% error rate, but if conversation completion = 20%, your agent isn't helping users. Technically healthy, practically useless.


3 levels of alerting

Alerts are the most common source of mistakes in production AI systems. Either there are too many (alert fatigue) or too few (you hear it from a customer).


Level 1: Info (a Slack channel, no paging)

What: daily summaries, weekly reports, general telemetry Channel: #ai-prod-info or an email digest Response time: whenever you get to it

Examples:

  • Daily at 09:00: "Yesterday: 1,247 requests, $12.40 spend, 0.3% error rate, p95 = 4.2s"
  • Weekly on Monday: "This week: cost trend +12%, top error: tool timeout (43 events)"
  • Monthly: "Monthly pricing optimization report"

Goal: context and trends, not reactive monitoring.


Level 2: Warning (a Slack mention, response within 1 hour)

What: degradation signals, problems that could become critical Channel: #ai-prod-alerts + @here Response time: 1 hour

Examples:

  • A cost spike 2x the daily average
  • Error rate >2% over 15 minutes
  • Latency p95 >2x the SLA
  • Cache hit ratio dropped below 30%
  • A single user with >100 requests in an hour (potential abuse)

Tool: a Slack incoming webhook + a structured message

python
# Example of sending a warning
def send_warning(metric, value, threshold):
    slack_webhook = os.getenv("SLACK_WARNING_WEBHOOK")
    payload = {
        "text": f"@here Warning: {metric} = {value} (threshold: {threshold})",
        "attachments": [{
            "color": "warning",
            "fields": [
                {"title": "Metric", "value": metric, "short": True},
                {"title": "Current", "value": str(value), "short": True},
                {"title": "Threshold", "value": str(threshold), "short": True},
                {"title": "Dashboard", "value": "https://link-to-your-dashboard", "short": True}
            ]
        }]
    }
    requests.post(slack_webhook, json=payload)

Level 3: Critical (page on-call, response within 5 min)

What: customer-impacting issues, the system is down, security incidents Channel: PagerDuty / a phone call / SMS Response time: 5 minutes

Examples:

  • The system is down (>50% errors)
  • A cost spike >5x the daily average (a potential breach or an infinite loop)
  • A security alert (an unusual access pattern, a suspected leaked key)
  • A customer-reported outage confirmed by the metrics
  • A critical SLA breach

Tool:

🎨 Picture this: the three levels are like a fire alarm system. A wisp of smoke in the kitchen (Info): good to know, no panic. A burning smell in the hallway (Warning): check it out and put it out. Flames on the floor (Critical): call 911 and evacuate. If every wisp of smoke wakes up the fire department, they'll show up to the real fire tired and grumpy.


4 observability tools in 2026: a comparison

Tool Free plan Best for
LangSmith Yes; see the website for limits Advanced tracing, prompt comparison
Helicone Yes; since March 2026 the service is in maintenance mode (acquired by Mintlify), no new features LLM-specific, cost tracking, quick setup
Sentry Yes; see the website for limits Error tracking, not LLM-specific
PostHog A monthly allowance for product analytics (see the website for the current size) User analytics, not traces

Paid plan prices change for all four; check the services' websites.

Recommendations by stage:

  • Personal / starter (<1K users): the free plan of LangSmith or PostHog plus the Anthropic Console is enough for the first few months
  • Team (1K-10K users): LangSmith + Sentry
  • Production (there's revenue): LangSmith + Sentry + PostHog + an on-call tool (PagerDuty or similar)
  • Enterprise: Datadog (https://www.datadoghq.com/) + custom dashboards

Also:


An implementation pattern: a proxy, using Helicone as the example

One of the fastest ways to get observability is to route the Anthropic SDK through a proxy, such as Helicone. Setup takes 5 minutes.

⚠️ As of October 2026, Helicone is in maintenance mode (see above), so treat it as an example of the approach: a proxy between your code and the API, plus tags in the headers. For a new project, compare it with LangSmith (the MLOps for indies lesson) and with gateways like Cloudflare's AI Gateway. Check the proxy URL and header names against the documentation of whichever service you choose.

Before (no observability):

python
from anthropic import Anthropic

client = Anthropic()
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}]
)

After (wrapped with Helicone):

python
from anthropic import Anthropic
import os

client = Anthropic(
    base_url="https://anthropic.helicone.ai",  # Proxy through Helicone
    default_headers={
        "Helicone-Auth": f"Bearer {os.getenv('HELICONE_KEY')}",
        "Helicone-User-Id": user_id,                    # Per-user tracking
        "Helicone-Property-Feature": "chat-bot",        # Feature breakdown
        "Helicone-Property-Environment": "production",  # Env tagging
        "Helicone-Cache-Enabled": "true",               # Bonus: caching
        "Helicone-Property-Tier": user_tier             # Custom dimension
    }
)

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}]
)
# Now the dashboard shows: latency, cost, tokens, user, feature, environment

What shows up in the dashboard right away:

  • Every request with a trace ID
  • Cost per request with a breakdown
  • Latency p50/p95/p99
  • Top users by spend
  • Top features by usage
  • Cache hit ratio

Dashboard structure: what belongs on the main screen

A good production dashboard is one you look at once a day for 30 seconds and understand everything.

Code
┌──────────────────────────────────────────────────────────┐
│ My App Production Dashboard             [Last 24h ▼]     │
├──────────────────────────────────────────────────────────┤
│                                                           │
│  📊 LAST 24H OVERVIEW                                     │
│  ┌─────────┬─────────┬─────────┬─────────┐               │
│  │ 1,247   │ $12.40  │  0.3%   │  4.2s   │               │
│  │ requests│  spend  │ errors  │   p95   │               │
│  └─────────┴─────────┴─────────┴─────────┘               │
│                                                           │
│  📈 TREND (7 days)                                        │
│  Cost:    ▁▂▃▃▄▅▆  +12% wow                              │
│  Errors:  ▁▁▂▁▁▁▁  stable                                │
│  p95:     ▃▃▃▃▃▃▃  stable                                │
│                                                           │
│  🚨 ACTIVE ALERTS                                         │
│  • [WARN] Cache hit ratio 24% (target >30%)              │
│                                                           │
│  🐢 TOP ISSUES                                            │
│  1. Slowest endpoint: /api/research (p95 = 18s)          │
│  2. Most errors: tool_timeout (43 events)                │
│  3. Biggest spender: user_42 ($3.20 today, 25% of total) │
│                                                           │
│  😊 USER SATISFACTION                                     │
│  CSAT: 4.3 / 5.0  (n=87)                                 │
│  Escalation rate: 12%                                     │
│                                                           │
└──────────────────────────────────────────────────────────┘

Rules for a good dashboard:

  • The main numbers at the top, on one screen, no scrolling
  • Trends shown visually (sparklines), not as tables
  • Active alerts visible immediately, in red
  • Top issues hint at where to dig
  • A business metric (CSAT) on equal footing with the technical ones

The debugging workflow: from complaint to fix

Scenario: a user writes, "The AI gave a wrong answer at 2:30 p.m."

Step 1: Locate: find the traces around 2:30 p.m. for this user in the logs

Type this into the chat
# In the tracing dashboard: filter user_id = user_42, time 14:25 — 14:35

Step 2: Reproduce: what was the input, what was the context, what was the output

  • Take the exact prompt from the trace
  • Run it in the Anthropic Console playground with the same parameters
  • Get the same (or a different) result

Step 3: Correlate: with the system metrics at that time

  • Was there a latency spike? (slow = less time to think)
  • Was there an error spike in the API? (degraded service)
  • Was there a cost spike? (a loop bug during that period)
  • Was there a cache miss? (a cold cache → different behavior)

Step 4: Hypothesize the cause: the options:

  • Context overflow (>180K tokens)? → a compaction strategy
  • Bad data (corrupted input)? → input validation
  • A bug in the system (a race condition)? → a unit test
  • LLM degradation (a model update)? → check the eval suite
  • Prompt drift (something changed)? → version-control your prompts

Step 5: Fix + prevent

  • A minimal fix
  • Add a test case to the eval suite (see the Evals: self-improving skills lesson)
  • Add an alert rule if it's a recurring pattern
  • Document it in the runbook

🎨 Picture this: debugging is a medical diagnosis. Symptom (the complaint) → history (traces) → lab tests (metrics) → diagnosis (the cause) → treatment (the fix) + recommendations (prevention). Without observability, you're a doctor without lab tests, guessing.


Logging best practices

✅ What to log:

  • All LLM requests/responses (input, output, tokens, latency, cost)
  • All tool calls (name, args, result, duration)
  • All errors with the full stack trace
  • User actions (without PII: pseudonymize)
  • System events (deploys, config changes, restarts)

❌ What NOT to log:

  • Plain-text passwords / API keys
  • Full PII (emails → hash, names → "User_42")
  • Credit card numbers
  • Full health data content
  • Internal IPs / system paths (security)

Sampling strategy:

  • 100% of errors (always)
  • 5-10% of normal traces (cost balance)
  • 100% of slow requests (>2x p95)
  • 100% of high-cost requests (>$0.50)

Retention policy:

Log type Keep for
Errors 90 days
Normal traces 30 days
Sampled traces (for trends) 1 year
Audit logs (compliance) 7 years (an example: the period depends on the country and industry; check with a lawyer)

Structured logging, always JSON:

python
import json
import time

def log_llm_call(request, response, duration_ms):
    log_entry = {
        "ts": time.time(),
        "level": "INFO",
        "event": "llm_call",
        "trace_id": request.headers.get("X-Trace-ID"),
        "user_id_hash": hash_user_id(request.user_id),
        "model": response.model,
        "tokens_input": response.usage.input_tokens,
        "tokens_output": response.usage.output_tokens,
        "cache_read_tokens": response.usage.cache_read_input_tokens or 0,
        "duration_ms": duration_ms,
        "cost_usd": calculate_cost(response.usage, response.model),
        "feature": request.feature_tag,
        "success": True
    }
    print(json.dumps(log_entry))  # → stdout → log aggregator

Anti-patterns (what not to do)

❌ Adding logging "later": usually after something has already broken. By then, there's no data to debug yesterday's incident.

❌ An alert for every error: noise, alert fatigue. Within a week the team mutes the channel. Filter by severity and frequency.

❌ Error tracking only, no cost tracking: you wake up to a surprise $5,000 bill. In the early stages, cost monitoring is more critical than error monitoring.

❌ Logging without a user_id: you can't debug user reports. "It doesn't work for me" → you can't find their trace.

❌ "It's all in console.log": not persistent, not searchable, not correlated. Logs should live in an aggregator (Helicone / Datadog / a homemade one).

❌ 100% sampling all the time: you pay to store everything. Sample wisely: errors 100%, normal 5%.

❌ Alerts without a runbook: an alert fires and the on-call person doesn't know what to do. Every critical alert should link to a runbook.

❌ A dashboard nobody looks at: if you open it once a month, you find out about problems a month late. The daily ritual matters.


The cost of observability

The most common question is "how much does this cost?" The honest answer:

Category Cost
Tools (starter) Free plans are often enough (LangSmith, Sentry, PostHog)
Tools (production) Paid plans from several services: work it out from their current prices
Setup time (initial) 4-8 hours
Maintenance 1-2 hours/month (updating alerts, tuning thresholds)
Storage (if self-hosted) Depends on log volume and the storage plan (S3, R2)

The ROI math: one missed cost spike (an infinite loop of calls, for example) can cost more than all your monitoring tools for a year. One missed customer-facing outage usually costs even more.

🎨 Picture this: observability is car insurance. When nothing's happening, it seems like a waste. One accident pays back ten years of premiums. Except that, unlike insurance, you can prevent the accident by seeing the problem coming.


Recommendations by maturity level

Beginner (no revenue, learning):

  • Just the Anthropic Console (built in) + manual checks weekly
  • No extra tools
  • Logs to the console / a file
  • Goal: understand what's going on at all

Intermediate (early users, free plans):

  • The free plan of LangSmith or PostHog
  • A Slack webhook for critical alerts
  • A weekly dashboard review
  • Goal: catch problems before users do

Professional (production with revenue):

  • LangSmith (paid plan) + Sentry + PostHog
  • PagerDuty or similar for critical alerts
  • Custom dashboards with business metrics
  • A daily review ritual
  • Goal: SLA compliance, proactive optimization

Enterprise (scale, compliance):

  • Datadog or New Relic full stack
  • A custom OpenTelemetry pipeline
  • Multi-region observability
  • Goal: never hear about problems from customers at all

A quarterly observability review

Once a quarter, set aside 2 hours to review your observability setup.

What to add:

  • New endpoints / features → new metrics
  • New SLOs from clients → new alerts
  • New types of errors have appeared → categorize and track them

What to remove:

  • Unused dashboards (nobody has looked at them in 90 days)
  • False-positive alerts (they fire, but it's not critical)
  • Metrics that have never been used to make a decision

What to optimize:

  • The noisiest alerts → tune the thresholds
  • Opportunities to reverse cost trends (where you can save)
  • Slow queries / endpoints

What to rotate:

  • Access keys (API tokens for observability tools)
  • The sub-processor list (compliance)
  • Runbook URLs (if something moved)

Practice

Step 1: Set up Helicone (5 minutes)

Helicone is used here as an example of the proxy approach: since March 2026 it's been in maintenance mode. If you don't want to depend on a service like that, replace this step with the LangSmith integration from the MLOps for indies lesson.

bash
# 1. Sign up at helicone.ai (there's a free plan)
# 2. Get an API key from the dashboard
# 3. Save it in .env

echo "HELICONE_API_KEY=sk-helicone-xxx" >> .env
echo "ANTHROPIC_API_KEY=sk-ant-xxx" >> .env
python
# helicone_setup.py — a minimal integration
import os
from dotenv import load_dotenv
from anthropic import Anthropic

load_dotenv()

# Client through the Helicone proxy
client = Anthropic(
    api_key=os.getenv("ANTHROPIC_API_KEY"),
    base_url="https://anthropic.helicone.ai",
    default_headers={
        "Helicone-Auth": f"Bearer {os.getenv('HELICONE_API_KEY')}",
        "Helicone-Property-Environment": "production",
        "Helicone-Property-App": "my-agent"
    }
)

# A test request
response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=100,
    messages=[{"role": "user", "content": "Say hello"}],
    extra_headers={
        "Helicone-User-Id": "test_user_001",
        "Helicone-Property-Feature": "greeting"
    }
)

print("".join(b.text for b in response.content if b.type == "text"))
print(f"\nCheck the dashboard: https://www.helicone.ai/dashboard")
print(f"Cost: visible in real time")

Step 2: Set up Slack alerts (15 minutes)

python
# alerts.py — three-level alerting
import os
import requests
from enum import Enum

class AlertLevel(Enum):
    INFO = "good"        # green
    WARNING = "warning"  # yellow
    CRITICAL = "danger"  # red

def send_alert(level: AlertLevel, title: str, message: str, dashboard_url: str = None):
    """Send an alert to Slack based on severity level."""

    webhook_map = {
        AlertLevel.INFO: os.getenv("SLACK_INFO_WEBHOOK"),
        AlertLevel.WARNING: os.getenv("SLACK_WARNING_WEBHOOK"),
        AlertLevel.CRITICAL: os.getenv("SLACK_CRITICAL_WEBHOOK"),
    }

    webhook = webhook_map[level]
    mention = "@here" if level == AlertLevel.WARNING else ("@channel" if level == AlertLevel.CRITICAL else "")

    payload = {
        "text": f"{mention} *{title}*",
        "attachments": [{
            "color": level.value,
            "text": message,
            "fields": [
                {"title": "Severity", "value": level.name, "short": True},
                {"title": "Dashboard", "value": dashboard_url or "N/A", "short": True}
            ]
        }]
    }

    response = requests.post(webhook, json=payload)
    response.raise_for_status()

# Usage examples
send_alert(
    AlertLevel.INFO,
    "Daily Summary",
    "Yesterday: 1247 requests, $12.40 spend, 0.3% errors"
)

send_alert(
    AlertLevel.WARNING,
    "Cost spike detected",
    "Already $40 today (2x yesterday's $20). Check the logs for the last hour.",
    "https://link-to-your-dashboard"
)

send_alert(
    AlertLevel.CRITICAL,
    "System degradation",
    "Error rate 12% over the last 5 minutes. p99 latency 30s.",
    "https://link-to-your-dashboard"
)

Step 3: A cost monitoring daemon

python
# cost_monitor.py — checking for cost anomalies
import os
import time
from datetime import datetime, timedelta
from anthropic import Anthropic

# Say you have a get_daily_cost() function that reads from your tracing service's API
def get_daily_cost(date):
    """Returns the $ spent in a day. A stub: connect your tracing service's API."""
    # Real implementation: a request to the tracing service's API for the given date
    return 12.40  # placeholder

def get_baseline(days_back=7):
    """Average cost over the last N days."""
    costs = []
    for i in range(1, days_back + 1):
        date = datetime.now() - timedelta(days=i)
        costs.append(get_daily_cost(date))
    return sum(costs) / len(costs)

def check_cost_anomaly():
    today_cost = get_daily_cost(datetime.now())
    baseline = get_baseline(days_back=7)

    ratio = today_cost / baseline if baseline > 0 else 0

    if ratio > 5.0:
        send_alert(
            AlertLevel.CRITICAL,
            "🚨 Cost spike >5x baseline",
            f"Today: ${today_cost:.2f}, baseline: ${baseline:.2f} ({ratio:.1f}x)"
        )
    elif ratio > 2.0:
        send_alert(
            AlertLevel.WARNING,
            "⚠️ Cost spike 2x baseline",
            f"Today: ${today_cost:.2f}, baseline: ${baseline:.2f} ({ratio:.1f}x)"
        )

# Runs hourly via cron
if __name__ == "__main__":
    check_cost_anomaly()
bash
# crontab -e
# Check for cost anomalies every hour
0 * * * * cd /path/to/project && python cost_monitor.py

# Daily summary at 09:00
0 9 * * * cd /path/to/project && python daily_summary.py

Step 4: Structured logging

python
# structured_logging.py — JSON logs ready for aggregation
import json
import time
import hashlib
from contextlib import contextmanager

def hash_user_id(user_id: str) -> str:
    """Pseudonymize user_id for logs."""
    return hashlib.sha256(user_id.encode()).hexdigest()[:16]

@contextmanager
def trace_llm_call(user_id: str, feature: str, model: str):
    """A context manager that logs an LLM call as structured JSON."""
    trace_id = f"trace_{int(time.time()*1000)}"
    start = time.time()

    log_data = {
        "ts": time.time(),
        "trace_id": trace_id,
        "user_id_hash": hash_user_id(user_id),
        "feature": feature,
        "model": model,
    }

    try:
        yield log_data
        log_data["success"] = True
    except Exception as e:
        log_data["success"] = False
        log_data["error"] = str(e)
        log_data["error_type"] = type(e).__name__
        raise
    finally:
        log_data["duration_ms"] = int((time.time() - start) * 1000)
        print(json.dumps(log_data))  # → stdout → log aggregator

# Usage
with trace_llm_call(user_id="user_42", feature="chat", model="claude-sonnet-5-5") as log:
    response = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello"}]
    )
    log["tokens_input"] = response.usage.input_tokens
    log["tokens_output"] = response.usage.output_tokens
    # Prices per 1M tokens as of October 2026 (Sonnet 5.5); current ones: ../actual.html
    PRICE_IN, PRICE_OUT = 2.0, 10.0
    log["cost_usd"] = response.usage.input_tokens * PRICE_IN / 1_000_000 + \
                     response.usage.output_tokens * PRICE_OUT / 1_000_000

Step 5: Build your own dashboard (optional)

python
# simple_dashboard.py — a minimal HTML dashboard
from flask import Flask, render_template_string
import time

flask_app = Flask(__name__)

DASHBOARD_TEMPLATE = """
<!DOCTYPE html>
<html>
<head>
    <title>AI Agent Dashboard</title>
    <meta http-equiv="refresh" content="60">
    <style>
        body { font-family: monospace; background: #1a1a1a; color: #00ff00; padding: 20px; }
        .metric { display: inline-block; padding: 20px; border: 1px solid #00ff00; margin: 10px; }
        .big { font-size: 32px; }
        .alert { color: #ff0000; }
        .warn { color: #ffaa00; }
        .ok { color: #00ff00; }
    </style>
</head>
<body>
    <h1>📊 AI Agent Production Dashboard</h1>
    <p>Last refresh: {{ ts }}</p>

    <div>
        <div class="metric">
            <div>Requests (24h)</div>
            <div class="big">{{ requests }}</div>
        </div>
        <div class="metric">
            <div>Cost (24h)</div>
            <div class="big ok">${{ cost }}</div>
        </div>
        <div class="metric">
            <div>Error rate</div>
            <div class="big {{ 'alert' if error_rate > 5 else 'warn' if error_rate > 1 else 'ok' }}">{{ error_rate }}%</div>
        </div>
        <div class="metric">
            <div>p95 latency</div>
            <div class="big">{{ p95 }}s</div>
        </div>
    </div>

    {% if alerts %}
    <h2>🚨 Active alerts</h2>
    <ul>{% for a in alerts %}<li class="warn">{{ a }}</li>{% endfor %}</ul>
    {% endif %}
</body>
</html>
"""

# Register the route with add_url_rule so the dashboard renders the data
def render_dashboard():
    # Plug in real data from the Helicone API / your own DB
    return render_template_string(
        DASHBOARD_TEMPLATE,
        ts=time.strftime("%Y-%m-%d %H:%M:%S"),
        requests=1247,
        cost="12.40",
        error_rate=0.3,
        p95=4.2,
        alerts=["Cache hit ratio 24% (target >30%)"]
    )

flask_app.add_url_rule("/", "dashboard", render_dashboard)

if __name__ == "__main__":
    flask_app.run(port=8080)
bash
# Open http://localhost:8080 and you'll see the dashboard
python simple_dashboard.py

Production readiness checklist (✅)

Before you let your agent loose on real users:

If you have fewer than 7 checkmarks, you're not ready for production. Finish the job.


Tools and resources

  • LangSmith: observability from LangChain, advanced tracing, with a free plan
  • Helicone: an LLM proxy with metrics (in maintenance mode as of October 2026)
  • Sentry: error tracking, general-purpose
  • PostHog: product analytics, business metrics
  • PagerDuty: on-call alerts, the industry standard
  • Datadog: enterprise full-stack observability
  • Grafana Cloud IRM: on-call and incidents (the open-source version of OnCall was archived in March 2026)
  • Anthropic Console: built-in usage tracking
  • OpenTelemetry: a vendor-neutral observability standard

Key takeaways

Observability isn't something for "later, once we've grown." It's what makes running in production possible at all. One missed cost spike can cost more than a year's subscription to a monitoring service. Set up something, anything: even a simple Slack webhook beats "I'll hear about it from customers."

5 mandatory metrics: latency (p95!), cost (spike detection), error rate (categorized), token usage (cache ratio), business (conversation completion). Missing any of the five leaves a blind spot that will bite you.

3 levels of alerting protect you from alert fatigue. Info: a daily summary. Warning: a Slack mention. Critical: page on-call. If every error wakes someone up at night, within a week the team mutes the channel and misses the real incident.

The debugging workflow is always the same: locate the trace → reproduce → correlate with metrics → hypothesize → fix + add an eval test. Without observability, you guess. With it, you diagnose.


Next lesson

→ The final architecture: a business's complete AI stack

For what to do when everything breaks, see the Backup & disaster recovery for your AI stack lesson.

The mark stays in this browser only and is never sent anywhere. My progress