The gist
Claude Code is like electricity from the grid. You pay a subscription and the lights come on (current prices and versions: What's current). It's convenient and reliable, but the power only flows while the bill is paid and the internet is up.
By October 2026 there are plenty of mature frameworks that let you run agent systems entirely on your own computer. No API. No internet. With tool use, memory, MCP servers and multi-agent collaboration.
The lesson Local AI models: Ollama, LM Studio and private AI covered basic local models (Ollama, LM Studio): those were the "light bulbs." This lesson is about a full "power plant": when a model + an agent framework + tools work together and solve real tasks without sending a single byte outside.
We'll go through 8 key frameworks, examples of models for agent tasks, the hardware requirements, and an honest bottom line: where local wins and where it loses to an API.
🎯 Decision tree: when local beats an API
The main question isn't "which is cooler," it's "which fits your task":
→ Production B2B with an SLA, support, uptime guarantees?
→ API (Anthropic/OpenAI). Local isn't built for this.
→ Working with confidential data (healthcare, legal, corporate code)?
→ LOCAL, required. Lawyers and doctors don't send client data outside.
→ Personal automation without the internet (a home bot, notes, translations)?
→ LOCAL. No point paying for private tasks.
→ Hobby projects and experiments?
→ LOCAL. Free and no limits.
→ High volume + simple tasks (classification, translating thousands of records)?
→ LOCAL is cheaper. An API will eat your budget fast.
→ Complex reasoning, ADRs, architecture decisions, complex code?
→ API (strong cloud models). Local isn't there yet.
→ Access to cloud AI APIs is limited where you or your client are, or you can't depend on a vendor's policies?
→ LOCAL = independence. No one can switch it off. (Not every country is on the supported-countries lists of Anthropic and OpenAI.)
→ Not sure whether you need maximum performance or savings?
→ Hybrid. Local for routine work, an API for the hard stuff.Key concepts
- Local agent framework: software that turns a local model (Llama, Qwen, Mistral) into a full agent with tools, memory and planning
- Function calling: a model's ability not just to answer in text but to call tools with the right arguments. A basic ability of agent models
- MCP (Model Context Protocol): an open protocol for connecting tools to models. It started at Anthropic and is now supported by local frameworks
- Inference backend: the engine that runs the model: Ollama (simple), vLLM (fast, production), SGLang (max throughput), llama.cpp (minimal dependencies)
- Quantization: compressing a model to run on smaller hardware. Q4_K_M lets a 70B model fit in 40GB instead of 140GB
- Open weights: the model is published with its weights (you can download and run it). Not the same as open source (open training code)
- Tool-native model: a model trained to call functions out of the box. Modern families (Qwen3, gpt-oss, Hermes 4) are tool-native. Older base models without a separate fine-tune often aren't
- Hybrid stack: a local + API combination: routine work locally, hard tasks through an API
- MoE (Mixture of Experts): an architecture where only some of the parameters are active for each request. For example, DeepSeek-V3 has 671B parameters with 37B active; Qwen3-Coder 30B activates about 3.3B. Big-model quality at mid-size speed
Theory
When local agents beat an API
Local isn't a replacement for an API. It's a different tool with different economics. Three situations where local really wins:
1. Privacy and sovereignty. Medical records, corporate code, personal notes: this is data that shouldn't leave your computer. APIs and subscriptions have different terms on training with your data, and a client's lawyers may say "no" regardless. Local solves this at the architecture level: the data physically never leaves.
2. Cost at high volume. If you have tens of thousands of text classification tasks a month, your API bill grows with the volume, while a local model in the 7–14B class costs only electricity. Do the math with your own volumes: token prices are on the What's current page.
3. Independence. Sanctions, outages, policy changes. For builders in some countries this isn't a hypothetical risk: not every country is on Anthropic's and OpenAI's supported-countries lists, and access to their APIs is limited there. Local always works.
When local loses: complex reasoning (the strongest cloud models), high-end multimodal, long context (hundreds of thousands of tokens and up), the newest features (computer use, voice). Cloud models are usually ahead here.
Hardware requirements: what you actually need
The main limit of local AI is hardware. The size of the model determines the minimum RAM/VRAM:
| Scenario | Mac | Windows/Linux PC | Budget |
|---|---|---|---|
| Starter (7–8B models) | A Mac with an M-series chip, 16 GB of memory | a graphics card with 12 GB of memory or more | minimal |
| Standard (13–14B + tools) | A Pro/Max-class Mac, 32 GB | a 16 GB graphics card | medium |
| Advanced (30–34B + agent stack) | A Max-class Mac, 64 GB | a 24 GB graphics card or two with 12–16 GB | high |
| Professional (70B+) | Mac Studio, 128–192 GB | a 48 GB graphics card or two with 24 GB | very high |
Hardware prices fluctuate a lot, so the table has no dollar amounts: check with sellers before you buy. The specific graphics card and chip models change every year; what matters is the amount of memory.
Important rules:
- A 70B-parameter model with Q4 quantization needs ~40GB of RAM/VRAM. On 32GB it'll run slowly (swap), and on 16GB it won't run at all
- The Mac advantage: unified memory. An M2 Max with 64GB behaves like 64GB of VRAM. On a PC, RAM and GPU VRAM are separate
- Older GPUs (GTX 1080, RTX 2060) work, but slowly. For a real-time agent you need a graphics card with at least 12 GB of memory
- Electricity: a powerful graphics card at full load draws 300–450 W. At 8 hours a day, that's a noticeable addition to your bill; calculate it with your own rate
8 key local agent frameworks (as of October 2026)
Let's go through them in order. Each has its own niche.
Framework 1: nanoClaude: a minimal agent for learning
What: Minimal Claude Code-style agent implementations that fit in a small amount of code. It's not one project but a whole genre: GitHub has several independent community projects (for example, nanoclaude and nano-claude-code). The idea is close to Karpathy's nanoGPT. Don't confuse it with NanoClaw: that's a different project, a lightweight containerized alternative to OpenClaw.
Best for: Education. Understanding how an agent works from the inside, without framework wrappers.
Runs on: Usually Python. They connect to a cloud model or to a local one through Ollama (see the specific project's README).
Cost: $0.
Status: Community-built; quality and maintenance vary. Not for production: it's a learning tool.
Source: https://github.com/karpathy/nanoGPT (the idea), https://github.com/CohleM/nanoclaude (an example of a learning agent).
Framework 2: OpenClaw: an open-source personal agent
What: An open-source, self-hosted personal AI agent (MIT license). It lives on your computer, talks to you through messaging apps (WhatsApp, Telegram, Discord, Slack and others), remembers past conversations, and can work with your email, calendar, browser and files. It works with both cloud models (Claude, GPT) and local ones. The project is run by the nonprofit OpenClaw Foundation, and there's no paid version.
Best for: People who want a personal agent under their own control and don't want to hand their data to someone else's service. For coding work, Claude Code, Codex and IDE agents are a better fit than OpenClaw.
Runs on: Mac, Windows, Linux. Installed with a script, through npm, or as a desktop app. The backend is any model, including a local one through Ollama.
Cost: $0 (open source); you only pay for the cloud model you choose, if you use one.
Status: Actively developed; check the official website for the current version.
Source: https://openclaw.ai and https://github.com/openclaw/openclaw.
Framework 3: Hermes: Nous Research's models and agent
What: Under the Hermes name, Nous Research releases two related things. First, the Hermes family of fine-tuned models (the current line is Hermes 4), trained on function calling and agent tasks. Second, Hermes Agent, an open-source agent framework (MIT license) with persistent memory, messaging app support and delegation to subagents. You can connect the agent to your own model or to Nous Portal.
Models: Hermes 4 on Hugging Face: 14B, 36B (version 4.3), 70B and 405B (405B needs a server GPU).
Best for: Function calling without an API, multi-step reasoning offline. When you need a model that calls tools correctly out of the box.
Runs on: Ollama / vLLM / TGI. Any framework that supports models of the corresponding family.
Cost: $0 (models and agent) + electricity for compute.
Source: https://huggingface.co/NousResearch and https://hermes-agent.nousresearch.com.
Framework 4: Qwen-Agent: Alibaba's native framework
What: An agent framework built by the Qwen team for its own models. Tight integration with Qwen models (Qwen3 and newer). Supports function calling, MCP, a code interpreter and RAG.
Models: Qwen3 (from 0.6B to 235B parameters) and Qwen3-Coder (30B and 480B). Check the project page for newer Qwen generations.
Best for: Complex coding tasks locally, multi-step workflows. If the task is "write and test a function," Qwen3-Coder is the first one to try.
Special: The framework itself is under the Apache 2.0 license, and many Qwen models also have open weights and Apache 2.0; check the specific model's license.
Runs on: Ollama / SGLang / vLLM. Native support.
Cost: $0.
Source: https://github.com/QwenLM/Qwen-Agent.
Framework 5: DeepSeek: a reasoning model offline
What: Open models from the Chinese lab DeepSeek. DeepSeek-R1 (early 2025) was one of the first reasoning models with open weights. As of October 2026, the current line is DeepSeek V4: V4.1-Flash came out on September 10, 2026, and its weights are on Hugging Face.
Models: V4.1-Flash is a large MoE model; as a rule you can't run it on ordinary home hardware and need a server. For home hardware, look at the smaller distilled versions of R1 (7B, 14B, 32B, 70B) and small models from other families.
Best for: Reasoning tasks offline. Math problems, complex code logic, multi-step planning.
Special: Open weights. There's also a paid API with low prices (current prices: What's current).
Cost: $0 (local, if your hardware is enough) or at the API rate.
Source: https://github.com/deepseek-ai.
Framework 6: Open WebUI + tools: a full local ChatGPT stack
What: A ChatGPT-like web UI for local models. The backend is Ollama. Supports tools, RAG, web search through SearxNG (local), a code interpreter, voice (TTS/STT through Whisper), and connecting MCP servers.
Best for: A chat interface on your own computer. When you need a "local ChatGPT" with advanced features.
Features: Multi-user (your family or team can use it), RAG (upload documents and the model can see them), pipelines (custom logic). One of the most popular projects in this niche. Check the license in the repository: it has terms about keeping the branding.
Runs on: Docker (starts with one command).
Cost: $0.
Source: https://github.com/open-webui/open-webui.
Framework 7: smolagents: Hugging Face's minimal framework
What: A minimalist agent framework from Hugging Face. The main idea: the agent writes code instead of calling tools through JSON schemas. Less overhead, more capability.
Best for: Experimentation, fast prototyping, education. When you want to see "what if the agent writes its own Python instead of using a limited set of tools."
Runs on: Any model (local through Ollama / API through providers).
Cost: $0 (framework).
Source: https://github.com/huggingface/smolagents.
Framework 8: AutoGen Studio + local models
What: AutoGen is a Microsoft framework for multi-agent collaboration. AutoGen Studio is a GUI on top of it. It works with local models through an OpenAI-compatible API. Important: as of October 2026, AutoGen is in maintenance mode, there won't be new features, and its successor is the Microsoft Agent Framework. AutoGen Studio is fine for prototypes, but not for production.
Best for: Multi-agent collaboration offline. When you need 3+ agents that "talk" to each other to solve a task.
Runs on: Ollama / LM Studio / any OpenAI-compatible endpoint. AutoGen Studio provides a visual builder.
Cost: $0 (framework).
Source: https://github.com/microsoft/autogen.
Also frequently mentioned:
- Continue.dev: an IDE agent for VS Code and JetBrains. Works with local models. Open source. An alternative for people who don't want a SaaS editor
- LM Studio: a GUI for managing local models + basic agent features. Good for beginners
- Jan.ai: an open-source ChatGPT alternative. Fully local. A simple UI
- CrewAI local: regular CrewAI with an Ollama backend instead of OpenAI
Examples of local models for agent tasks (as of October 2026)
Models are the heart of a local agent. Without the right model, even the best framework is useless:
| Model | Size | For | Tool calling | Download size in Ollama |
|---|---|---|---|---|
| Qwen3-Coder | 30B MoE (~3.3B active) | Coding tasks | ✅ | see the model page |
| Qwen3 | 8B / 14B / 30B / 32B | General-purpose, fast agents | ✅ | 5.2 / 9.3 / 19 / 20 GB |
| gpt-oss (OpenAI, Apache 2.0) | 20B / 120B | General + reasoning with adjustable effort | ✅ | 14 GB (needs 16 GB of RAM or more) / 65 GB (80 GB GPU) |
| Hermes 4 | 14B / 36B / 70B / 405B | Function calling, agent tasks | ✅ | depends on size and quantization |
| DeepSeek V4.1-Flash | large MoE | Reasoning, coding on a server | see the documentation | needs a server |
The Ollama download sizes were checked against the model pages as of October 2026. How to choose: don't trust percentages like "almost as good as Sonnet" from other people's comparisons. Run your own task on two or three models (see Step 5) and compare the results yourself. On complex architecture tasks, the gap with strong cloud models is bigger.
Practical picks:
- Mac with 32GB → Qwen3 14B or gpt-oss 20B (fits comfortably)
- Mac with 64GB+ or a 24GB graphics card → Qwen3 30B / Qwen3-Coder 30B / Qwen3 32B (the sweet spot for quality/speed)
- Mac Studio with 128GB+ or a server → gpt-oss 120B, Hermes 4 70B and larger models
Setup walkthrough: a quick start in 30 minutes
A full local agent stack takes one evening to set up:
# Step 1: Install Ollama (Mac/Linux)
curl -fsSL https://ollama.com/install.sh | sh
# Windows: download the installer from ollama.com
# Step 2: Download a model (once, ~19GB)
ollama pull qwen3:30b
# Alternatives by size:
# ollama pull qwen3:14b # ~9GB, for 16GB RAM
# ollama pull qwen3:8b # ~5GB, the minimum
# Current names and sizes: ollama.com/library
# Step 3: Install Open WebUI with Docker
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:main
# Step 4: Open http://localhost:3000 — your local ChatGPT is readyAfter setup, you have:
- A local ChatGPT-like interface
- A model that covers a good share of everyday tasks (code, writing, translation). Check the quality on your own tasks
- Tool calling (the model can call functions)
- RAG (upload a PDF and the model answers based on its contents)
- Full privacy (not a byte leaves your machine)
Real-world cases: what actually works locally
Not theory: patterns that already work:
Case 1: A privacy-first translator (a lawyer)
- Stack: Qwen3 14B + Open WebUI
- Use: translating clients' confidential documents EN↔︎ES
- Speed: depends on the hardware; measure it on your own computer
- Cost: no token charges after the initial setup
- Why local: client confidentiality rules out third-party services like DeepL or the Claude API
Case 2: Code review for a proprietary codebase
- Stack: Qwen3-Coder + Continue.dev in VS Code
- Use: code review without sending proprietary code to the Anthropic API
- Quality: good enough for standard reviews; complex architecture reviews are better left to a strong cloud model
- Cost: no token charges (you need a Mac with 64GB or a PC with a 24GB graphics card)
- Why local: the company's IP never leaves, and the NDA isn't violated
Case 3: Personal automation (a home bot)
- Stack: a small model (Qwen3 8B or Hermes 4 14B) + Open WebUI + custom tools
- Use: smart home (Home Assistant integration), personal notes, reminders
- Privacy: 100% offline
- Hardware: a compact Mac or a mini PC with 16GB of memory (one-time)
- Cost: no monthly fee
Case 4: A local agent for content (a newsletter)
- Stack: Qwen3 14B + smolagents + RAG over a local Obsidian vault
- Use: writing posts in the style of your notes and previous posts
- Cost: no token charges after setup
- Trade-off: the quality is lower than strong cloud models. For a personal newsletter, it's fine
- Why local: your drafts and ideas stay with you
Limits and trade-offs: honestly
No rose-colored glasses. Where local is genuinely weaker:
| What an API does better | What local does better |
|---|---|
| Reasoning quality (the strongest cloud models) | Privacy + sovereignty |
| Speed on simple queries | Cost at high volume |
| The newest features (computer use, voice) | Offline capability |
| No setup hassle | Customization (fine-tuning, prompts) |
| High-end multimodal (vision, audio) | Independence from a vendor |
| Long context (hundreds of thousands to a million tokens) | Predictable (fixed) cost |
| Tool ecosystem (MCP, plugins, marketplace) | Hardware = a one-time cost |
| Speed on small tasks | No rate limits |
The honest bottom line as of October 2026: for most professional tasks, an API still wins. Local is for specific use cases (privacy / cost / offline / sovereignty). But the gap is narrowing: the latest generations of open models (Qwen3, gpt-oss, DeepSeek V4) cover many everyday tasks. The strongest cloud models are still ahead on complex reasoning and long tasks.
Setup cost: one-time vs. ongoing
The most common question is "when does it pay for itself":
| Setup | One-time | Monthly ongoing |
|---|---|---|
| Minimum (Mac with 16GB + Ollama + a 7–8B model) | $0 (if you already have the Mac) | $0 + a small amount of electricity |
| Standard (Mac with 32GB + a 14–30B model + Open WebUI) | $0 for software + the price of a new Mac | $0 + electricity |
| Professional (PC with a 24GB+ graphics card + a 70B model + full stack) | the cost of building the PC | $0 + electricity (noticeably more because of the GPU) |
Compared with an API:
- A subscription to a cloud assistant costs twelve months' worth per year (current prices: What's current)
- The payback formula: hardware cost ÷ (monthly API bill − electricity). Hardware only pays for itself at high task volume, or when you need privacy
- Local doesn't replace an API 100%. In practice, builders use a hybrid: local for routine work, an API for the hard stuff
- Hybrid economics: local takes the high-volume simple tasks, and the cloud model's bill covers only the hard ones
Anti-patterns: what NOT to do
The rakes beginners step on:
❌ Using local 7B models for production B2B. The quality isn't there. The client will notice. For production, use strong cloud models, period.
❌ Running a 70B model on 8GB of RAM. It'll grind to a halt in swap. The minimum for 70B is 40GB of unified memory or 48GB of VRAM.
❌ Ignoring electricity costs. A GPU at full load for many hours a day raises the bill noticeably. In some scenarios, this eats up the savings vs. an API.
❌ Trying to replicate the strongest cloud models locally. The gap is real. Don't torture your hardware.
❌ Skipping security. Local ≠ automatically safer. A model from an unverified Hugging Face repository may carry malicious code in its scripts. Download from trusted sources, and don't enable trust_remote_code without reading the code.
❌ One local instance for the whole team without queue management. A bottleneck. If 5 people send requests at the same time, the model will stall. You need a load balancer (vLLM) or separate instances.
❌ Comparing local with an API on one task and drawing conclusions. Local wins on volume and privacy. An API wins on complex reasoning. Compare on a specific use case.
By audience: who should use what
A progression by level:
Beginner (just getting started):
- Stack: Ollama + Qwen3 14B + Open WebUI
- Covers: most personal tasks (translation, writing/editing, research)
- Hardware: a computer with 16GB of memory
Intermediate (has a coding background):
- Stack: + smolagents + Continue.dev in the IDE
- Covers: an agent-style local workflow for development
- Hardware: a Mac with 32GB or a PC with a 16GB graphics card
Professional (builder, agency):
- Stack: + AutoGen or the Microsoft Agent Framework / Hermes 4 70B + custom offline MCP servers
- Covers: a full multi-agent local stack for proprietary projects
- Hardware: a Mac Studio with 64GB+ or a PC with a 24GB graphics card
Enterprise (team):
- Stack: vLLM server + load balancer + fine-tuned models + auth
- Covers: an internal AI platform for the whole team
- Hardware: a dedicated server with several professional GPUs
Trends: where local AI is heading
What becomes possible every 6 months:
Smaller, faster. What was frontier a couple of years ago now runs on a laptop. Models of the same quality keep getting smaller.
Tool-native by default. New models (Qwen3, gpt-oss) are trained with function calling out of the box, so separate fine-tunes like Hermes are needed less often.
Multimodal local. Local models can work with images (for example, Qwen3-VL). Audio (local Whisper) is standard. Video is catching up.
Reasoning offline. DeepSeek-R1 showed that reasoning capability is possible with open weights. Today gpt-oss, Qwen3 and DeepSeek all have reasoning modes, and they increasingly fit on consumer hardware.
Mobile AI. Small models (1–4B) already run on smartphones. The next step is agents on mobile without the cloud.
MCP standardization. MCP has become the standard for connecting tools. Local frameworks (Open WebUI, smolagents, Qwen-Agent) support MCP servers, so the same tools work with the Claude API and with local models.
Practice
Step 1: Install Ollama and your first model
# Mac/Linux, with one command
curl -fsSL https://ollama.com/install.sh | sh
# Check the installation
ollama --version
# It should show a version number
# Download Qwen3 14B (a good balance for 32GB of RAM)
ollama pull qwen3:14b
# If you have less memory, start with 8B
ollama pull qwen3:8b
# Test the model
ollama run qwen3:14b "Hi, how are you?"
# It should answer in EnglishStep 2: Test function calling with Python
# test_function_calling.py — checking that the model can call tools
import ollama
# Describe the tool as a function
def get_weather(city: str) -> str:
"""Returns the weather for a city (a stub)"""
weather_db = {
"Chicago": "23°F, snow",
"Cuenca": "64°F, sunny",
"city_x": "59°F, fog"
}
return weather_db.get(city, "No data")
# The tool description for the model
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for the given city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name"
}
},
"required": ["city"]
}
}
}]
# Send the request
response = ollama.chat(
model='qwen3:14b',
messages=[{
'role': 'user',
'content': "What's the weather in Cuenca?"
}],
tools=tools
)
# Look at the result
print(response['message'])
# It should contain tool_calls with a call to get_weather(city="Cuenca")
# Simulate running the tool
if response['message'].get('tool_calls'):
for tool_call in response['message']['tool_calls']:
if tool_call['function']['name'] == 'get_weather':
city = tool_call['function']['arguments']['city']
result = get_weather(city)
print(f"Tool result: {result}")Step 3: Install Open WebUI with Docker
# Make sure Docker is installed
docker --version
# Start Open WebUI
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:main
# Wait ~30 seconds, then open it
open http://localhost:3000
# On first launch:
# 1. Create an admin account (local, it doesn't go anywhere)
# 2. In Settings → Connections → Ollama should be picked up automatically
# 3. In Chat, select the qwen3:14b model
# 4. Start a conversationStep 4: A simple local agent with smolagents
# local_agent.py — a minimal agent with smolagents and Ollama
from smolagents import CodeAgent, OpenAIModel, tool
# Connect to Ollama through the OpenAI-compatible endpoint
# (in older versions of smolagents this class was called OpenAIServerModel)
model = OpenAIModel(
model_id="qwen3:14b",
api_base="http://localhost:11434/v1",
api_key="ollama" # Ollama ignores the key, but requires a string
)
# Create a custom tool
@tool
def calculate_compound_interest(principal: float, rate: float, years: int) -> float:
"""
Calculate compound interest.
Args:
principal: the starting amount
rate: the annual rate (for example, 0.05 for 5%)
years: the number of years
"""
return principal * ((1 + rate) ** years)
@tool
def read_file(path: str) -> str:
"""
Read a file from disk.
Args:
path: the path to the file
"""
with open(path, 'r', encoding='utf-8') as f:
return f.read()
# Create the agent
agent = CodeAgent(
tools=[calculate_compound_interest, read_file],
model=model,
add_base_tools=True # Adds python_interpreter and others
)
# Run it
result = agent.run(
"If I put $1000 in at 7% annual interest for 10 years, "
"how much will I have at the end? Explain the calculation."
)
print("\n" + "="*50)
print("RESULT:")
print("="*50)
print(result)# Installation
pip install smolagents
# Run
python local_agent.py
# The agent should:
# 1. Understand the task
# 2. Call calculate_compound_interest(1000, 0.07, 10)
# 3. Get ~1967.15
# 4. Explain the calculation in EnglishStep 5: Compare local vs. API on your own task
# compare_local_vs_api.py — an honest comparison on one task
import time
import ollama
from anthropic import Anthropic # pip install anthropic
# Your task: pick something real
prompt = """Write a Python function that:
1. Takes a list of numbers
2. Returns the top 3 largest values
3. Includes type hints
4. Has a docstring with an example
5. Handles edge cases (empty list, < 3 elements)"""
# Option 1: Local through Ollama
print("=" * 50)
print("LOCAL: Qwen3 14B")
print("=" * 50)
start = time.time()
local_response = ollama.chat(
model='qwen3:14b',
messages=[{'role': 'user', 'content': prompt}]
)
local_time = time.time() - start
print(local_response['message']['content'])
print(f"\nTime: {local_time:.1f}s | Cost: $0")
# Option 2: API through Anthropic
print("\n" + "=" * 50)
print("API: Claude Sonnet 5.5")
print("=" * 50)
client = Anthropic() # Requires ANTHROPIC_API_KEY in env
start = time.time()
api_response = client.messages.create(
model="claude-sonnet-5-5", # current model IDs: Anthropic's documentation
max_tokens=1000,
messages=[{"role": "user", "content": prompt}]
)
api_time = time.time() - start
input_tokens = api_response.usage.input_tokens
output_tokens = api_response.usage.output_tokens
# Prices per 1M tokens: Sonnet 5.5 is $2 input, $10 output (as of October 2026).
# Current prices: the What's current page on the Academy website
PRICE_IN, PRICE_OUT = 2, 10
cost = (input_tokens * PRICE_IN + output_tokens * PRICE_OUT) / 1_000_000
print("".join(b.text for b in api_response.content if b.type == "text"))
print(f"\nTime: {api_time:.1f}s | Cost: ${cost:.4f}")
# Summary
print("\n" + "=" * 50)
print("COMPARISON:")
print("=" * 50)
print(f"Local: {local_time:.1f}s, $0")
print(f"API: {api_time:.1f}s, ${cost:.4f}")
print(f"\nIf you run 1000 tasks like this a month:")
print(f"Local: $0 + electricity")
print(f"API: ${cost * 1000:.2f}")What this test will show:
- On simple coding tasks, local is often comparable in quality
- Local latency is usually higher than an API's
- At high task volume, local starts to win on cost (plug your own volume into the calculation)
- The quality of the final code: judge it yourself by reading both answers
Tools and resources
- Ollama: the simplest backend for local models. Mac/Linux/Windows
- Open WebUI: a ChatGPT-like UI for local models
- Hermes (Nous Research): the tool-native Hermes 4 models
- Hermes Agent: an open-source agent from Nous Research
- OpenClaw: an open-source personal agent in your messaging apps
- Qwen Team: the main page for Qwen models and frameworks
- Qwen-Agent: the agent framework for Qwen
- smolagents (Hugging Face): a minimal Python agent framework
- DeepSeek: DeepSeek models with open weights
- AutoGen: Microsoft's framework for multi-agent systems (maintenance mode; the successor is the Microsoft Agent Framework)
- Continue.dev: an open-source IDE agent
- LM Studio: a GUI for managing local models
- Jan.ai: an open-source local ChatGPT alternative
- vLLM: a production-grade inference engine
- SGLang: a max-throughput inference framework
- Karpathy nanoGPT: the educational foundation for nano agents
- Prices and versions: What's current
Key takeaways
Local agents aren't a replacement for an API: they're a different tool with different economics. Privacy, cost at volume, offline, independence: four reasons local really wins. For everything else, an API is still ahead.
Hardware is the bottleneck. A computer with 16–32GB of memory can run an 8–14B-class model and cover many personal tasks. A Mac Studio or a PC with a 24GB+ graphics card is for serious work with 30B models and up. Electricity counts too.
The realistic path is a hybrid stack. Local for routine work (classification, translation, simple coding, RAG over your personal documents), an API for the hard stuff (architecture, complex reasoning, multimodal). Don't pick one: combine them.
What's next
→ Self-hosted AI enterprise stack: vLLM, SGLang, Docker: how to build an internal AI platform for a team when one computer is no longer enough
The mark stays in this browser only and is never sent anywhere. My progress