Library · Reliability: monitoring, failures, backups

The final architecture: a business's complete AI stack

Builder95 minUpdated: October 2026
92 of 105 in the library

Time: about 35 min theory + 60 min practice


The gist

You've learned each tool on its own. Now you're looking at everything at once: how it all comes together into a working business. Architecture isn't a diagram that looks pretty on a slide. It's a map that tells you where to run when something breaks, and what to add when it's time to grow.

🎨 Picture this: you've studied the hammer, the saw, the level and the drill. This lesson is the first time you see the blueprint for the whole house. Suddenly it clicks: the hammer and the saw don't compete. They're used at different stages of the same build.


Key concepts

  • The 5-layer stack: Frontend → API → AI → Data → Payments
  • The integration layer: n8n/Zapier as the glue between tools
  • Microsoft integration: Teams, Office, Power Platform for corporate clients
  • Stack cost: what it's made of at each level: Starter / Pro / Scale
  • Growth points: where to expand once the money shows up

Theory

The complete architecture of an AI business in 2026

Code
┌─────────────────────────────────────────────────────────┐
│                       USER                               │
└────────────────────────┬────────────────────────────────┘
                         │
┌────────────────────────▼────────────────────────────────┐
│               FRONTEND (Layer 1)                        │
│  v0.app / Webflow / Next.js / Telegram Bot              │
│  What's here: UI, forms, dashboards, mobile layout      │
└────────────────────────┬────────────────────────────────┘
                         │
┌────────────────────────▼────────────────────────────────┐
│               API GATEWAY (Layer 2)                     │
│  Cloudflare Workers / FastAPI / Express                 │
│  What's here: routing, auth, rate limiting, logging     │
└─────────┬──────────────┬──────────────┬─────────────────┘
          │              │              │
┌─────────▼──────┐ ┌─────▼──────┐ ┌────▼──────────────────┐
│  AI (Layer 3)  │ │ DATA (L.4) │ │ PAYMENTS (Layer 5)    │
│  Claude API    │ │ Supabase   │ │ Stripe                │
│  Embeddings   │ │ PostgreSQL │ │ Subscriptions         │
│  Vector DB    │ │ Redis cache│ │ Webhooks              │
└────────────────┘ └────────────┘ └───────────────────────┘
          │
┌─────────▼────────────────────────────────────────────────┐
│           INTEGRATION LAYER (cross-cutting)              │
│  n8n / Zapier: connecting everything to everything      │
│  CRM → Marketing → Support → Analytics                  │
└──────────────────────────────────────────────────────────┘

Layer 1: Frontend, what the user sees

Options and when to use them:

Tool When Price
v0.app (formerly v0.dev) Quick prototype, React Has a free plan; prices on the website
Webflow Marketing site without code Paid; prices on the website
Next.js Full control, SEO matters Free (hosting is separate)
Telegram bot B2C, mass-market users Free
Streamlit Internal tool, MVP Free

🎨 Picture this: a store window. Shoppers see the window, not the warehouse behind it. Start with the simplest window that sells.

Layer 2: API Gateway, the nervous system

Cloudflare Workers are the best fit for indie builders:

javascript
// worker.js: basic structure
export default {
  async fetch(request, env) {
    const url = new URL(request.url);
    
    // Routing
    if (url.pathname === '/api/generate') {
      return handleGenerate(request, env);
    }
    
    if (url.pathname === '/api/webhook/stripe') {
      return handleStripeWebhook(request, env);
    }
    
    return new Response('Not found', { status: 404 });
  }
};

async function handleGenerate(request, env) {
  // Auth check
  const apiKey = request.headers.get('X-API-Key');
  if (!await validateKey(apiKey, env)) {
    return new Response('Unauthorized', { status: 401 });
  }
  
  // Rate limiting with KV
  const count = await env.RATE_LIMITS.get(apiKey) || 0;
  if (count > 100) {  // 100 requests per hour
    return new Response('Rate limit exceeded', { status: 429 });
  }
  await env.RATE_LIMITS.put(apiKey, String(parseInt(count) + 1), {expirationTtl: 3600});
  
  // Call Claude
  const body = await request.json();
  const result = await callClaude(body.prompt, env.ANTHROPIC_API_KEY);
  
  return Response.json({ result });
}

Layer 3: AI, the brain of the system

Claude as the central AI, but with the right routing between models:

python
def route_to_model(task_type: str, complexity: str) -> str:
    """
    Cut costs by routing each task to the right model.
    """
    # Prices per 1 million tokens (input / output) as of October 2026; current ones: ../actual.html
    routing = {
        ("classification", "simple"): "claude-haiku-4-5",      # $1 / $5
        ("summarization", "simple"): "claude-haiku-4-5",
        ("generation", "standard"): "claude-sonnet-5-5",        # $2 / $10
        ("generation", "complex"): "claude-sonnet-5-5",
        ("reasoning", "critical"): "claude-opus-5-5",           # $4 / $20
        ("architecture", "critical"): "claude-opus-5-5",
    }
    return routing.get((task_type, complexity), "claude-sonnet-5-5")

# Example
model = route_to_model("classification", "simple")
# → claude-haiku-4-5 (the lightest and cheapest in the lineup)

Layer 4: Data, the system's memory

Code
Supabase (PostgreSQL): main storage
├── users (accounts, settings)
├── conversations (chat history)  
├── documents (uploaded files)
└── analytics (events, metrics)

pgvector (Supabase's vector extension): semantic search
├── embeddings for RAG
└── finding similar documents

Redis (Upstash): fast cache
├── rate limiting
├── sessions
└── results of frequent requests

The integration layer: n8n as the conductor

n8n is no-code/low-code automation. It connects everything to everything:

Code
New customer in Stripe
        ↓ webhook
      n8n workflow:
        ├── Create a record in Supabase
        ├── Send a welcome email (Gmail/SendGrid)
        ├── Add to the CRM (HubSpot/Notion)
        ├── Notification to Slack or Telegram (for you)
        └── Start the onboarding sequence

When to use n8n and when to use Zapier:

n8n Zapier
Self-hosted ✅ (free for internal use; check the license terms on the website) ❌
Cloud Starter €20/mo billed annually (as of October 2026) Free: 100 tasks a month; Professional from $19.99/mo billed annually
Custom code ✅ JS nodes ⚠️ limited
Complex workflows ✅ ⚠️
Ease of use ⚠️ ✅

For indie builders: n8n self-hosted on a VPS = automation without a subscription fee, but the server and updates are on you. VPS prices depend on the hosting provider.

Microsoft integration for corporate clients

If your clients are corporations, they run everything on Microsoft. Here's how to get in:

Teams + Claude via a webhook:

Microsoft is retiring the old Office 365 Connectors webhooks. You now create a webhook through the Workflows app in Teams (the "Send webhook alerts to a channel" template), and the message carries an Adaptive Card:

python
import os
import requests

def notify_teams_channel(webhook_url: str, message: str, facts: dict = None):
    """Send a notification to a Microsoft Teams channel via a Workflows webhook (Adaptive Card)."""
    body = [{"type": "TextBlock", "text": message, "weight": "Bolder", "wrap": True}]
    if facts:
        body.append({
            "type": "FactSet",
            "facts": [{"title": k, "value": v} for k, v in facts.items()]
        })

    payload = {
        "type": "message",
        "attachments": [{
            "contentType": "application/vnd.microsoft.card.adaptive",
            "content": {
                "$schema": "http://adaptivecards.io/schemas/adaptive-card.json",
                "type": "AdaptiveCard",
                "version": "1.2",
                "body": body
            }
        }]
    }

    response = requests.post(webhook_url, json=payload)
    return response.status_code in (200, 202)

# Example: notification that an AI analysis is complete
notify_teams_channel(
    webhook_url=os.environ["TEAMS_WEBHOOK_URL"],
    message="✅ AI analysis of the financial reports is complete",
    facts={"Documents processed": "47", "Time": "3.2 min", "Anomalies found": "2"}
)

SharePoint as a knowledge base for Claude:

python
import os
import anthropic
import requests
from msal import ConfidentialClientApplication

def get_sharepoint_documents(site_id: str, folder: str) -> list:
    """Get documents from SharePoint via the Graph API."""
    app = ConfidentialClientApplication(
        client_id=os.environ["AZURE_CLIENT_ID"],
        client_credential=os.environ["AZURE_CLIENT_SECRET"],
        authority=f"https://login.microsoftonline.com/{os.environ['AZURE_TENANT_ID']}"
    )
    
    token = app.acquire_token_for_client(
        scopes=["https://graph.microsoft.com/.default"]
    )
    
    headers = {"Authorization": f"Bearer {token['access_token']}"}
    
    # Get the files in the folder
    url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drives/root:/{folder}:/children"
    response = requests.get(url, headers=headers)
    
    files = response.json().get("value", [])
    return [{"name": f["name"], "url": f["@microsoft.graph.downloadUrl"]} 
            for f in files if f.get("file")]

def answer_from_sharepoint(question: str, site_id: str) -> str:
    """Claude answers a question using SharePoint documents."""
    docs = get_sharepoint_documents(site_id, "Knowledge Base")
    
    # Load the document contents
    context = ""
    for doc in docs[:5]:  # Top 5 most recent
        content = requests.get(doc["url"]).text[:2000]
        context += f"\n---{doc['name']}---\n{content}"
    
    client = anthropic.Anthropic()
    return "".join(b.text for b in client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1000,
        messages=[{
            "role": "user",
            "content": f"Company documents:\n{context}\n\nQuestion: {question}"
        }]
    ).content if b.type == "text")

OpenAI models on Azure vs Claude for corporations:

OpenAI models on Azure Claude
Where it runs Inside Azure In Microsoft Foundry: on Azure infrastructure or on Anthropic infrastructure (two hosting options)
Regions and data zone See the Azure documentation Depend on the model and hosting option: check the Microsoft Foundry documentation
GDPR and data processing Check the contract and documentation Check the contract and documentation for the option you choose
Azure integration Native Through Microsoft Foundry, with setup and billing through Azure Marketplace
Model quality Depends on the task Depends on the task: compare both on your own data
Price Azure pricing Claude pricing in Foundry; see Microsoft's documentation
For corporations on Azure Often the default choice Depends on requirements

If a client says "we can't go outside Azure," suggest Microsoft Foundry (formerly Azure AI Foundry): Claude is available there in an option that runs entirely on Azure infrastructure. Setup goes through Azure Marketplace and requires a paid Azure subscription in a supported country. For the list of available models and regions, see the Microsoft Learn documentation.

The real cost of the stack at each level

Prices below are as of October 2026, and only where they've been checked against the services' websites. Work out the rest yourself from current pricing, and see the What's current page for Claude models.

Starter: your first clients

Component Tool Price
Frontend Cloudflare Pages Free + Next.js (the free Vercel Hobby plan is only for non-commercial projects) $0
API Cloudflare Workers Free (100,000 requests a day) $0
AI Claude API Depends on traffic and model
Database Supabase Free (2 projects, 500 MB database, paused after a week of inactivity) $0
Payments Stripe 2.9% + 30¢ in the US; different in other countries
Automation n8n self-hosted VPS price depends on the hosting provider
Email Resend Free Free plan; limits on the website
Monitoring LangSmith Free Free plan; limits on the website
Total Sum of the rows; at low traffic the biggest item is the Claude API

Pro: a growing business

Component Tool Price
Frontend Vercel Pro $20/mo per developer
API Cloudflare Workers Paid from $5/mo
AI Claude API (more traffic) Depends on traffic and model
Database Supabase Pro $25/mo
Vector DB Supabase pgvector included
Cache Upstash Redis Upstash pricing
Automation n8n Cloud Starter €20/mo billed annually
Email Resend Resend pricing
Monitoring LangSmith LangSmith pricing
Total Sum of the rows at current prices

Scale: a serious business

  • Claude API: the biggest expense at high traffic (Opus for critical tasks)
  • Supabase: plans above Pro; see the website
  • Dedicated infrastructure: at your hosting provider's rates
  • Teams + support tools: at the rates of the services you choose
  • Total: calculate it for your own stack

🎨 Picture this: Starter is a bicycle. Fast, cheap, and it gets you where you need to go. Pro is a car. Scale is a truck. Most people's mistake: they want the truck right away. Buy the bicycle, and prove you need the car.


Practice

Step 1: Draw the architecture of your AI business

Grab a sheet of paper (or Miro/FigJam) and draw:

  1. Who the user is: where they come in from, what they want
  2. Frontend: what they see, where it runs
  3. API: what handles requests
  4. AI: where Claude sits, which prompts, which model
  5. Data: what's stored, where it comes from
  6. Payments: how you make money

There's no "right" answer. The point is to see the whole system.

Step 2: Calculate the cost of your stack

Fill in a table following the Starter/Pro/Scale examples. What's free for you right now? What becomes paid as you grow?

Step 3: Find the bottleneck

Look at your architecture and answer these questions:

  • What breaks first at 10x the load?
  • Where do you depend on a single provider with no fallback?
  • Which link gets the most expensive as you scale?

Step 4: Assignment

If you have an idea for an AI product, draw its full architecture. If not, take any product from the course (a support chatbot, an AI content tool, email automation) and break down its stack across all 5 layers with real prices.


Tools and resources

  • Diagrams: draw.io (free), Miro, Excalidraw
  • n8n self-hosted: n8n.io/docs/hosting
  • Microsoft Graph API: learn.microsoft.com/graph
  • Teams Webhooks: learn.microsoft.com/microsoftteams/platform/webhooks-and-connectors
  • Microsoft Foundry (formerly Azure AI Foundry): ai.azure.com
  • Supabase: supabase.com
  • Cloudflare Workers: workers.cloudflare.com
  • Upstash Redis: upstash.com

Key takeaways

Architecture isn't about technology, it's about decisions. Each layer answers a question: Frontend, "what does the user see?"; API, "who routes?"; AI, "who thinks?"; Data, "what gets remembered?"; Payments, "how does the business make money?". Start with the Starter stack: at the beginning, free plans are usually enough. Scale when the load and revenue justify it, not before.


Next lesson

→ Unit economics of the AI stack: the full financial math at 3 levels

The mark stays in this browser only and is never sent anywhere. My progress