Library · Reliability: monitoring, failures, backups

MLOps for indie builders: monitoring, drift and retraining without a DevOps team

Engineer65 minUpdated: October 2026
88 of 105 in the library

Time: ~25 min theory + 40 min practice


The gist

MLOps sounds like something from a big tech company with a hundred engineers. In practice, it's the answer to one question: "Is my AI working as well as it did a week ago?" For an indie developer, MLOps comes down to three things: keep an eye on cost, keep an eye on quality, and don't break what already works.

🎨 Picture this: You opened a coffee shop. You dialed in your espresso recipe once and that was that. But a month later the supplier changed the beans, the machine's pressure drifted, and the barista started tamping a little differently. The coffee is "about the same." Customers don't complain out loud. They just stop coming. MLOps is tasting your own coffee every day.


Key concepts

  • Prompt drift: a prompt gets worse over time without any change to the code
  • LangSmith: tracing and monitoring for your Claude calls
  • Promptfoo: regression testing for prompts
  • Cost monitoring: budget alerts and spending anomalies
  • Prompt versioning: Git for prompts, not just for code
  • Baseline metrics: quality metrics you compare against over time

Theory

What prompt drift is and why it happens

You wrote a prompt three months ago. The code hasn't changed. But suddenly you notice the answers have gotten longer, or less specific, or a bit different in tone. You didn't touch anything. So what happened?

Causes of prompt drift:

  1. The model changed. Your code uses an alias instead of a specific version, or the old model was retired from the API and you moved to a new one. The behavior shifts a little or a lot. Pin the model version in your code and keep an eye on the list of models and retirements: What's current.

  2. The context changed. Your users started writing differently, the questions changed, new patterns appeared that the prompt didn't account for.

  3. The agent chain shifted. One agent started producing a slightly different format → the next agent stopped parsing it correctly.

  4. The system prompt is out of date. You described a context that no longer applies.

🎨 Picture this: Your GPS is running on last year's maps. The road is the same, but a new shopping center has blocked your usual route. The GPS guides you confidently, and you end up staring at a fence.

LangSmith: tracing your Claude calls

LangSmith (from LangChain) is an observability platform for AI. It records every call: the input prompt, the response, latency, cost and tags.

Why you need it:

  • See what's actually happening inside your agent chain
  • Compare answers "now vs. a week ago"
  • Find expensive calls you can optimize
  • Debug when something breaks in production

Integration with the Anthropic SDK:

python
import anthropic
from langsmith import traceable
from langsmith.wrappers import wrap_anthropic

# wrap_anthropic records the tokens and cost of every call
client = wrap_anthropic(anthropic.Anthropic())

@traceable(name="content-generator")
def generate_content(topic: str, style: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1000,
        messages=[{
            "role": "user",
            "content": f"Write content about '{topic}' in a {style} style"
        }]
    )
    return "".join(b.text for b in response.content if b.type == "text")

# Every call automatically shows up in the LangSmith dashboard
result = generate_content("email marketing", "friendly")

What you see in the dashboard:

  • Every call with the full prompt and response
  • Latency percentiles (p50, p90, p99)
  • Cost per endpoint
  • Anomalies (a call took 10 times longer than usual)

Promptfoo: testing your prompts

Promptfoo is a command-line tool for running regression tests on prompts. It works like unit tests, but for AI. In March 2026, Promptfoo was acquired by OpenAI; the company says the project stays open source, but keep an eye on how it develops.

bash
npm install -g promptfoo

Test configuration (promptfooconfig.yaml):

yaml
prompts:
  - "Analyze this review and identify the sentiment: {{review}}"

providers:
  - anthropic:messages:claude-haiku-4-5

tests:
  - vars:
      review: "Great product, very happy with it!"
    assert:
      - type: contains
        value: "positive"
      - type: not-contains
        value: "negative"
  
  - vars:
      review: "Terrible quality, never buying again"
    assert:
      - type: contains
        value: "negative"
      - type: llm-rubric
        value: "The answer should identify negative sentiment"
  
  - vars:
      review: "It's fine, nothing special"
    assert:
      - type: contains-any
        value: ["neutral", "mixed", "unclear"]
bash
# Run the tests
promptfoo eval

# Compare two versions of a prompt
promptfoo eval --prompts v1-prompt.txt v2-prompt.txt

CI/CD integration:

yaml
# .github/workflows/prompt-tests.yml
name: Prompt Regression Tests

on: [push, pull_request]

jobs:
  test-prompts:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - run: npm install -g promptfoo
      - run: promptfoo eval --no-progress-bar
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

Now the tests run on every push. If the prompt has regressed, promptfoo exits with an error and CI fails.

Cost monitoring: keeping your money under control

For an indie builder, money matters most. One bug in a prompt can multiply your daily spending.

Built-in cost monitoring:

python
import anthropic
from datetime import datetime, timezone
import json
import os

class CostTracker:
    def __init__(self, daily_budget_usd: float = 20.0):
        self.client = anthropic.Anthropic()
        self.daily_budget = daily_budget_usd
        self.today_cost = 0.0
        self.log_file = "cost_log.jsonl"
    
    # Prices per 1M tokens as of October 2026. They change: see the current ones on
    # the "What's current" page (../actual.html). Add other models the same way.
    COSTS = {
        "claude-sonnet-5-5": {"input": 2.0, "output": 10.0},
        "claude-haiku-4-5": {"input": 1.0, "output": 5.0},
    }
    
    def track_call(self, model: str, input_tokens: int, output_tokens: int, endpoint: str):
        rates = self.COSTS.get(model, self.COSTS["claude-sonnet-5-5"])
        cost = (input_tokens / 1_000_000 * rates["input"] + 
                output_tokens / 1_000_000 * rates["output"])
        
        self.today_cost += cost
        
        entry = {
            "ts": datetime.now(timezone.utc).isoformat(),
            "model": model,
            "endpoint": endpoint,
            "input_tokens": input_tokens,
            "output_tokens": output_tokens,
            "cost_usd": round(cost, 6)
        }
        
        with open(self.log_file, "a") as f:
            f.write(json.dumps(entry) + "\n")
        
        # Alert when 80% of the budget is used
        if self.today_cost > self.daily_budget * 0.8:
            self.send_alert(f"⚠️ Spent ${self.today_cost:.2f} of the ${self.daily_budget} budget")
        
        return cost
    
    def send_alert(self, message: str):
        # Notification through an incoming webhook (for example, Slack or Discord)
        import requests
        requests.post(
            os.environ['ALERT_WEBHOOK_URL'],
            json={"text": message}
        )

tracker = CostTracker(daily_budget_usd=20.0)

Spending anomalies:

python
def detect_cost_anomaly(log_file: str, window_days: int = 7):
    """Compares today's spending with the average for the past week."""
    import statistics
    
    with open(log_file) as f:
        entries = [json.loads(line) for line in f]
    
    today = datetime.now(timezone.utc).date()
    daily_costs = {}
    
    for entry in entries:
        date = datetime.fromisoformat(entry['ts']).date()
        daily_costs[date] = daily_costs.get(date, 0) + entry['cost_usd']
    
    recent = [cost for date, cost in daily_costs.items() 
              if (today - date).days <= window_days and date != today]
    
    if len(recent) < 3:
        return False
    
    avg = statistics.mean(recent)
    std = statistics.stdev(recent)
    today_cost = daily_costs.get(today, 0)
    
    # Anomaly: today costs more than 2 standard deviations above the average
    if today_cost > avg + 2 * std:
        return f"🚨 Anomaly! Today: ${today_cost:.2f}, average: ${avg:.2f}"
    
    return False

Versioning prompts like code

Prompts are code. They belong in Git with a history of changes.

Repository structure:

Code
prompts/
  v1/
    system-prompt.md     # Version 1.0
    user-template.md
  v2/
    system-prompt.md     # Version 2.0: changed the tone
    user-template.md
  current -> v2/         # Symlink to the current version
  CHANGELOG.md           # What changed and why

CHANGELOG.md for prompts:

markdown
## v2.0 — 2026-04-15

### Changes
- Removed the "be brief" instruction: answers were too short for business emails
- Added formatting examples
- Changed the tone from formal to "friendly-professional"

### Metrics before/after (promptfoo)
- Average answer length: 150 → 280 words ✅
- Test "contains a greeting": 60% → 95% ✅
- Test "no 'To whom it may concern'": 100% (unchanged) ✅

### Rollback
git checkout v1/ if v2 turns out worse

Practice

Step 1: Install the tools

bash
# Promptfoo for testing prompts
npm install -g promptfoo

# LangSmith (for tracing)
pip install langsmith

# Environment variables
export LANGSMITH_TRACING="true"   # tracing won't turn on without this
export LANGSMITH_API_KEY="ls__xxx"
export LANGSMITH_PROJECT="my-ai-app"

Step 2: Write your first prompt tests

Create a promptfooconfig.yaml file for your main prompt. Come up with 5-10 test cases, including edge cases: empty input, very long text, text in different languages, a potentially malicious request.

bash
promptfoo eval
# See what passed and what failed

Step 3: Set up cost tracking

python
# cost_tracker.py
# Copy the code from the lesson and set ALERT_WEBHOOK_URL
# Add a tracker.track_call() call after every messages.create()

Step 4: Basic monitoring

Run this once a day:

bash
# Daily total
python -c "
import json
from datetime import datetime, timezone

today = datetime.now(timezone.utc).date().isoformat()
total = 0
calls = 0

with open('cost_log.jsonl') as f:
    for line in f:
        e = json.loads(line)
        if e['ts'].startswith(today):
            total += e['cost_usd']
            calls += 1

print(f'Today: {calls} calls, \${total:.4f}')
"

Step 5: Assignment

  1. Take any prompt from your project
  2. Write 10 tests in promptfoo
  3. Run the tests and record your baseline (how many passed)
  4. Make a small change to the prompt
  5. Run them again: did it get better or worse?
  6. Commit both versions to Git with a description of the changes

Tools and resources

  • LangSmith: smith.langchain.com (there's a free plan; check the site for limits)
  • Promptfoo: promptfoo.dev, npm install -g promptfoo
  • Helicone, an alternative to LangSmith: helicone.ai (in March 2026 the company was acquired by Mintlify; the service is in maintenance mode with no new features, so for a new project it's better to pick another tool)
  • Braintrust, another evaluation and monitoring platform: braintrust.dev
  • Weights & Biases, enterprise MLOps: wandb.ai

Key takeaways

MLOps for indie builders isn't DevOps. It's discipline. Three rules: test your prompts before you deploy (Promptfoo), track spending with alerts (cost tracker), and log everything in LangSmith so you have somewhere to look when something goes wrong. Prompt drift quietly erodes quality. Without tests, you won't notice it until clients start leaving.


Next lesson

→ Production observability: what to monitor when your agent is live

The mark stays in this browser only and is never sent anywhere. My progress