Library · AI on your own computer and server

Local AI models: Ollama, LM Studio and private AI

Builder75 minUpdated: October 2026
76 of 105 in the library

Module: 13. Professional practice | Time: about 30 min theory + 45 min practice


The gist

Everything you tell Claude goes to Anthropic's cloud. For most tasks that's fine. But what if the client is a hospital, a law firm or a bank? What if you need to process 100,000 documents with no API budget? What if there's no internet?

Local models are AI that runs right on your computer. The data doesn't go anywhere. Inference costs next to nothing (electricity and wear on the hardware). It works on a plane.

About model names. The catalog of open models gets updated every few months. The examples in this lesson come from families that were popular when it was written (Llama 3, Mistral, Qwen 2.5, Phi, Gemma, DeepSeek). By October 2026 newer generations of these families had come out, along with other open models such as OpenAI's gpt-oss. The principle for choosing doesn't change: look up current names and sizes in the ollama.com/library catalog. Swap the command you need in for the tag in the example.

🎨 Picture this: cloud AI is like a cab. Convenient and reliable, but the driver sees and hears everything, and every ride costs money. Local AI is like your own car in the garage. You go wherever you want, whenever you want, and nobody's listening. You only pay for gas (electricity).


Key concepts

  • Ollama: the main tool for running local models (macOS, Windows, Linux)
  • LM Studio: a GUI alternative with a drag & drop interface
  • GGUF: a quantization format: a big model gets compressed to a reasonable size
  • Quantization: reducing the precision of the weights (32-bit → 4-bit) for speed and memory
  • OpenAI-compatible API: Ollama serves the same API format as OpenAI, so code written for the OpenAI API works with Ollama almost unchanged
  • Apple Silicon advantage: M-series chips use memory shared by the CPU and GPU, which is a big advantage for local models
  • LLM Router: a pattern: pick the model for the task instead of one model for everything

Theory

Why local models: five real reasons

Not theoretical reasons, but specific situations from practice.

Reason 1: Privacy because the client requires it

A medical clinic wants to automate processing patient records. A law firm wants to analyze confidential contracts. A bank wants to process internal documents.

They're not against AI. They're against patient/client/partner data going to Anthropic's or OpenAI's servers. This isn't paranoia: it's GDPR, HIPAA, NDAs.

A local model is the only way to give them AI without compliance risk.

Reason 2: Cost at high volume

Say you need to process 50,000 documents of 2 pages each. Through Claude Sonnet that's:

  • 50,000 × ~1,000 tokens = 50M tokens
  • Cost of the input tokens: 50 × the price per 1M tokens. At $2 per 1M (the Sonnet 5.5 price as of October 2026) that's about $100
  • Plus output tokens

If this repeats weekly, that's hundreds of dollars a month for this one task. For bulk background processing in the cloud there's the Batch API with a 50% discount, but even then the costs add up. And if the task doesn't need Sonnet-level quality (classification, data extraction), a local model will handle it almost for free. Current prices: What's current.

Reason 3: Offline, anywhere

A plane. A cabin in the woods. A construction site with no internet. A conference with bad Wi-Fi. An emergency when the cloud is down.

A local model always works. This isn't exotic: it's a real use case for anyone who works in the field.

Reason 4: Compliance regulations

Some countries and industries have strict data residency requirements: data can't leave a certain jurisdiction or server. Banks in a number of countries can't use public cloud AI at all.

For this client segment, local AI isn't an option, it's the only choice.

Reason 5: Fine-tuning on your own data

Want a model that speaks in your client's exact style? That knows their product by heart? That answers exactly the way their business needs?

You can't fine-tune Claude's weights (it's a closed model). But you can take an open model (for example, from the Llama, Qwen, Gemma or Mistral families), train it on a few hundred examples and get an assistant specialized for that specific business. More here: Fine-tuning: when prompts aren't enough.

🎨 Picture this: cloud models are like a universal Swiss Army knife. Local models with fine-tuning are like a tool made to fit your hand exactly, for one specific job.


Ollama: the main tool

Ollama is a runtime for local models. You install it once and launch any model from the catalog with a command. It downloads the model automatically, optimizes it for your hardware and starts an API server.

Installation:

bash
# macOS via Homebrew
brew install ollama

# Or download it directly: ollama.com/download
# Windows and Linux are available too, the installer is on the site

# Check the installation
ollama --version

Running your first model:

bash
# Run Llama 3.2 (3B parameters, ~2GB, fast on any Mac). The tag is an example: see the Ollama catalog for current models
ollama run llama3.2

# Now you have an interactive chat right in the terminal
>>> Hi! Explain what an API is in simple terms.

Popular models:

bash
# Fast (for simple tasks, run on 8GB RAM)
ollama run llama3.2          # 3B parameters, ~2GB of disk
ollama run phi4              # Microsoft, 14B, very efficient
ollama run gemma3            # Google, good quality for its size

# For code (optimized for code)
ollama run mistral:7b        # French, excellent general purpose
ollama run qwen2.5-coder     # Alibaba, specialized in code
ollama run deepseek-coder    # DeepSeek, strong at code and debugging
ollama run codellama         # Meta, built specifically for code

# Heavy (need 32GB+ RAM, Mac Studio level)
ollama run llama3.1:70b      # Noticeably stronger than small models
ollama run qwen2.5:72b       # Alibaba, very strong

# Management
ollama list                  # List downloaded models
ollama ps                    # What's running right now
ollama rm llama3.2           # Delete a model
ollama pull llama3.2         # Download without running

The key fact: an OpenAI-compatible API

Ollama starts a local API server at http://localhost:11434. This server is compatible with the OpenAI API format. That means code written for the OpenAI API usually works with Ollama almost unchanged: you just change the base_url and the model name.

bash
# The API runs automatically whenever Ollama is running
curl http://localhost:11434/v1/models

# Test with curl (just like OpenAI)
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [{"role": "user", "content": "Hi!"}]
  }'

LM Studio: a GUI for people who don't like the terminal

LM Studio (lmstudio.ai) is a visual interface for local models. If Ollama = a terminal tool, LM Studio = an app with an interface.

What it can do:

  • Download models from HuggingFace right inside the interface (search is built in)
  • Drag & drop installation of GGUF files
  • A built-in chat with history
  • A built-in API server (also OpenAI-compatible)
  • Parameter settings: temperature, context length, GPU layers
  • Real-time RAM and GPU monitoring

When LM Studio beats Ollama:

  • You're showing local AI to a client (it's visually easier to follow)
  • You need to try lots of different models quickly
  • You want to tweak parameters without the command line
  • You're working with GGUF files straight from HuggingFace

When Ollama beats LM Studio:

  • Automation with scripts
  • Server deployment (Linux, no GUI)
  • Integration into code (simpler through the CLI)
  • Quickly switching models in the terminal

🎨 Picture this: Ollama is like git on the command line. LM Studio is like GitHub Desktop. Same result, different interface. Pros use both.


Understanding parameters and RAM

Before you pick a model, you need to understand how parameters, size and quality relate.

What model parameters are

🎨 Picture this: parameters are like neurons in a brain. A human has ~86 billion neurons. A small 3B model has 3 billion weights. A big 70B model has 70 billion. More usually means smarter, but it needs more memory.

What quantization is

A model's original weights are stored at 32-bit precision (fp32). That's huge. Quantization compresses them to 4-bit (Q4) or 8-bit (Q8) precision with minimal loss of quality.

The result: a 7B model in Q4 takes ~4GB instead of ~28GB. It runs on a regular MacBook.

Type this into the chat
Q2: very heavy compression, noticeable loss of quality
Q4: a good balance (recommended for most)
Q5 / Q6: higher quality, more memory
Q8: nearly original quality, 2x the RAM of Q4
fp16: full precision, needs a professional GPU

A practical table: what runs on what

Computer RAM Largest model (Q4) What it gets you
8GB 7B parameters Good general tasks, basic code
16GB 13B parameters Noticeably better text quality
32GB 30B parameters Very good quality
64GB 70B parameters Strong previous-generation models
128GB+ 120B+ parameters Very large open models

Important for Apple Silicon: Apple's M-series chips use unified memory: the CPU and GPU share one pool of RAM. That's a big advantage. On a PC with a separate graphics card, the model has to fit in the card's VRAM (a card with 8GB of VRAM = models around 7B). On a Mac with 24GB, most of the memory is available to the model. The table above is approximate: the real size depends on quantization and context length.


Apple Silicon: why a Mac beats a PC for local AI

🎨 Picture this: a regular PC is like an office where the lawyers (CPU) and the accountants (GPU) work in different buildings and keep couriering documents back and forth. An Apple Silicon Mac is like an open-plan office where everyone sits together and passes documents instantly.

The architectural advantage:

On a PC:

  • The CPU has 64GB of RAM
  • The GPU has a separate 8GB of VRAM
  • A 70B model (needs 40GB) = doesn't fit in VRAM = runs slowly on the CPU

On a Mac M3/M4:

  • CPU + GPU = one memory pool
  • A Mac with 64GB of unified memory runs a 70B Q4 model
  • The GPU accelerates inference through the Metal framework

Speed: specific tokens/sec figures depend on the model, quantization, Ollama version and hardware, so there aren't any here. The general picture: a powerful graphics card is usually faster on small models, while a Mac with lots of unified memory wins on very large ones, because the big model simply fits. And a Mac is a laptop that goes everywhere with you. Measure the speed on your own hardware: ollama run <model> --verbose shows the eval rate in tokens per second.

A Mac with lots of memory is the tipping point:

If you're serious about local models, a Mac Studio with a very large amount of unified memory changes the picture: large models become practical for everyday chat.


Integrating with code: the LLM Router pattern

The most important pattern when working with local models is the LLM Router. Not one model for everything. Different tasks, different models.

🎨 Picture this: an LLM Router is like an HR director who knows who to give which task to. Complex analysis goes to Sonnet. Simple classification goes to the local model (free). Confidential data goes only to the local one (it doesn't leave the machine).

A simple router implementation:

python
# llm_router.py
import anthropic
import openai
from enum import Enum

class TaskType(Enum):
    COMPLEX_CODE = "complex_code"       # Claude Sonnet
    SIMPLE_TEXT = "simple_text"         # Local: llama3.2
    CODE_REVIEW = "code_review"         # Local: deepseek-coder
    PRIVATE_DATA = "private_data"       # Local: any (data doesn't leave)
    CREATIVE = "creative"               # Claude Sonnet (best quality)
    CLASSIFICATION = "classification"  # Local (cheap, fast)

def get_client_and_model(task_type: TaskType):
    """Returns (client, model_name) for the task"""
    
    # Anthropic Claude: complex tasks, creative work
    if task_type in [TaskType.COMPLEX_CODE, TaskType.CREATIVE]:
        return (
            anthropic.Anthropic(),
            "claude-sonnet-5-5"  # model name as of October 2026, check what's current
        )
    
    # Ollama: local, free
    ollama_client = openai.OpenAI(
        base_url="http://localhost:11434/v1",
        api_key="ollama"  # Ollama ignores the key, but the parameter is required
    )
    
    model_map = {
        TaskType.SIMPLE_TEXT: "llama3.2",
        TaskType.CODE_REVIEW: "deepseek-coder",
        TaskType.PRIVATE_DATA: "mistral:7b",  # Data doesn't leave the computer
        TaskType.CLASSIFICATION: "llama3.2",
    }
    
    return (ollama_client, model_map[task_type])


def run_task(task_type: TaskType, prompt: str) -> str:
    """Runs the task through the right model"""
    client, model = get_client_and_model(task_type)
    
    # Anthropic API (a different format)
    if isinstance(client, anthropic.Anthropic):
        message = client.messages.create(
            model=model,
            max_tokens=1024,
            messages=[{"role": "user", "content": prompt}]
        )
        return "".join(b.text for b in message.content if b.type == "text")
    
    # OpenAI-compatible API (Ollama)
    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}]
    )
    return response.choices[0].message.content


# Usage example
if __name__ == "__main__":
    # Complex code goes to Claude
    result = run_task(
        TaskType.COMPLEX_CODE,
        "Write an async Python function to process a Celery queue"
    )
    
    # Classification goes to the local model (free)
    category = run_task(
        TaskType.CLASSIFICATION,
        "Categorize this text: 'I'd like to know apartment prices in Cuenca'"
        "\nCategories: [price inquiry, neighborhood inquiry, general question]"
    )
    
    # Private data stays local only
    analysis = run_task(
        TaskType.PRIVATE_DATA,
        "Extract the key dates from this contract: [contract text]"
    )

A smarter router with automatic selection:

python
# smart_router.py
import re

ROUTING_RULES = [
    # (pattern in the prompt, task_type)
    (r"(confidential|nda|private|secret)", TaskType.PRIVATE_DATA),
    (r"(write code|function|class|async|debug)", TaskType.COMPLEX_CODE),
    (r"(categorize|classify|type|kind)", TaskType.CLASSIFICATION),
    (r"(check the code|code review|find bugs)", TaskType.CODE_REVIEW),
]

def auto_route(prompt: str) -> TaskType:
    """Automatically determines the task type from the prompt"""
    prompt_lower = prompt.lower()
    for pattern, task_type in ROUTING_RULES:
        if re.search(pattern, prompt_lower):
            return task_type
    return TaskType.SIMPLE_TEXT  # Default: the local model

# Usage
task = auto_route("Check this code for bugs: def foo(): pass")
result = run_task(task, "Check this code for bugs: def foo(): pass")

Comparing models for different tasks

The table reflects the model families at the time of writing; see the Ollama catalog for newer generations.

Model Parameters RAM (Q4) Best for Weaker at Speed
Llama 3.2 3B 3B 2GB Quick chat, simple questions Complex code, long texts Very fast
Phi-4 14B 9GB Reasoning, math Long context Fast
Mistral 7B 7B 5GB General tasks, following instructions Multi-step reasoning Fast
Qwen 2.5 Coder 7B 7B 5GB Code, SQL, debugging Creative writing Fast
DeepSeek Coder 6.7B 6.7B 4GB Code, explaining code Long context Fast
Llama 3.1 8B 8B 5GB Good general purpose No specialization Fast
Llama 3.1 70B 70B 40GB A strong previous-generation model Needs 64GB+ RAM Slow
Qwen 2.5 72B 72B 45GB Best local quality Needs 64GB+ RAM Slow

Recommendations for choosing:

Type this into the chat
Getting started (8-16GB RAM):
  General tasks: mistral:7b
  Code: qwen2.5-coder or deepseek-coder
  Quick and simple: llama3.2

Regular work (32GB):
  Main: llama3.1:8b (a good balance)
  Code: qwen2.5-coder:14b

Professional level (64GB+):
  Everything: llama3.1:70b
  Code: qwen2.5-coder:32b

Ollama MCP: local models right from Claude Code

You can add Ollama as an MCP server and call local models right from a Claude Code session. Useful when you want Claude to delegate part of a task to a local model.

bash
# Install the Ollama MCP (a package by a third-party author: check its repository before installing)
claude mcp add ollama-mcp -- npx -y ollama-mcp

# Check that it was added
claude mcp list

After installing it, you can write in a Claude Code session:

Type this into the chat
"Use ollama with the qwen2.5-coder model to check this code"

And Claude will delegate the task through MCP to your local Ollama. The data doesn't go to the cloud: the processing happens on your computer.

The alternative: a direct API call from Claude Code:

python
# Claude Code can write and run this script
import openai

client = openai.OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"
)

response = client.chat.completions.create(
    model="qwen2.5-coder",
    messages=[{
        "role": "user",
        "content": "Review this Python code for bugs:\n" + code_content
    }]
)
print(response.choices[0].message.content)

Local vs cloud: an honest comparison

You don't have to pick one. The pattern is to use both for different tasks.

Local models WIN at:

Type this into the chat
✅ Client data under NDA / HIPAA / GDPR
✅ Bulk document processing (thousands of them, no per-token charges)
✅ Working without internet
✅ Fine-tuning for a specific business
✅ Compliance in regulated industries
✅ Long R&D without a token meter running
✅ Repetitive simple tasks (classification, extraction)

The cloud (Claude Sonnet/Opus) WINS at:

Type this into the chat
✅ Complex multi-step tasks (better quality)
✅ Very large context (the big Claude models go up to a million tokens)
✅ Computer Use and browser automation
✅ A quick start with no 8GB download
✅ An always up-to-date model (it updates itself)
✅ Complex code, architectural decisions
✅ When you need the best result on the first try

🎨 Picture this: local AI is like having your own cook at home: cheap, private, available any time. Cloud Claude is like a Michelin-starred restaurant: for when you need the very best for an important occasion. A smart entrepreneur uses both depending on the situation.

Decision matrix:

Task Recommendation Why
A content draft that will be edited Local Cheap, it needs editing anyway
Final text for publication Cloud Quality is critical
Processing a client's 1,000 PDFs Local No per-token charges + the data doesn't leave
A complex agentic workflow Cloud Reliability and context
Classifying incoming requests Local Simple task, at volume
Generating code from scratch Cloud (Claude) Best quality
Code review of finished code Local (DeepSeek) Good enough for checking
Confidential analysis Local The data doesn't leave

Fine-tuning: when and how

Fine-tuning is further training of an already trained model on your own data. The result: a model that talks like your client, knows their product and answers in the right style.

When fine-tuning is worth it:

Type this into the chat
✅ A specific communication style (tone, brand terminology)
✅ Niche knowledge (medical terms, legal wording)
✅ A specific output format (always JSON with a certain structure)
✅ Bulk generation of the same kind of content
✅ When prompt engineering no longer gets you the quality you need

When you DON'T need fine-tuning:

Type this into the chat
❌ Just "I want ChatGPT to know more": that's RAG, not fine-tuning
❌ The task can be solved with a good prompt
❌ You have fewer than 200-300 examples
❌ No time for iteration (fine-tuning = experiments)

Tools for fine-tuning:

LoRA (Low-Rank Adaptation) is the most popular method. It doesn't retrain the whole model, it adds small adapters. Cheap in compute, fast.

Type this into the chat
Full fine-tuning of a 7B model: needs a serious cloud GPU (A100 class), rented by the hour
LoRA fine-tuning of a 7B model: a powerful consumer graphics card or a Mac with lots of memory is enough

Unsloth is a tool that makes LoRA 2-5x faster with lower memory use:

bash
pip install unsloth

# A basic fine-tuning example with Unsloth
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/llama-3.2-3b-instruct",
    max_seq_length=2048,
    load_in_4bit=True  # Quantization to save memory
)

Data format for fine-tuning:

jsonl
{"messages": [
  {"role": "user", "content": "How do I buy an apartment in Cuenca?"},
  {"role": "assistant", "content": "The buying process in Ecuador includes..."}
]}
{"messages": [
  {"role": "user", "content": "Do I need a notary for the purchase?"},
  {"role": "assistant", "content": "Yes, in Ecuador all real estate transactions..."}
]}

At least 200-300 pairs like these. Claude Code can help you generate the dataset: you describe the style and the typical questions, and it creates the examples.

Where to train:

Type this into the chat
Runpod.io: cloud GPUs billed by the hour (prices on the site)
Google Colab: cheaper, but with limits
A Mac with lots of unified memory: locally, if you have the hardware
Modal.com: serverless GPUs, you pay only for training time

Practice: install Ollama + write a document processor

Step 1: Installation and first run

bash
# Install
brew install ollama

# Start the service (in the background)
ollama serve &

# Download and run your first model
ollama run mistral:7b

Once it starts, you'll see an interactive chat. Type something and make sure it works. Ctrl+D to exit.

bash
# Check that the API works
curl http://localhost:11434/api/tags
# Should return a list of models

# For code, download one more model
ollama pull qwen2.5-coder

Step 2: A Python script for processing documents

Create a file called document_processor.py:

python
"""
A local document processor using Ollama.
The data doesn't leave your computer.
"""

import openai
import json
from pathlib import Path

# Connect to the local Ollama
client = openai.OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"
)

def extract_info_from_document(text: str, model: str = "mistral:7b") -> dict:
    """
    Extracts structured information from the text of a document.
    Uses a local model, so the data doesn't leave.
    """
    prompt = f"""Analyze this document and extract information in JSON format.

Find:
- the document date (field "date")
- the parties to the contract (field "parties", a list)
- the amount, if any (field "amount")
- the document type (field "type")
- the key obligations (field "obligations", a list)

Return ONLY valid JSON with no explanations.

DOCUMENT:
{text}"""

    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        temperature=0.1  # Low temperature for structured output
    )
    
    raw = response.choices[0].message.content.strip()
    
    # Clean up if the model added markdown
    if raw.startswith("```"):
        raw = raw.split("```")[1]
        if raw.startswith("json"):
            raw = raw[4:]
    
    try:
        return json.loads(raw)
    except json.JSONDecodeError:
        return {"error": "Could not parse the JSON", "raw": raw}


def classify_document(text: str, model: str = "llama3.2") -> str:
    """
    Quick classification of the document type.
    We use a small model: cheap and fast.
    """
    prompt = f"""Identify the document type with one word from this list:
[contract, invoice, receipt, letter, application, other]

DOCUMENT (first 500 characters):
{text[:500]}

Answer with a single word only."""

    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        max_tokens=10
    )
    return response.choices[0].message.content.strip().lower()


def process_folder(folder_path: str) -> list:
    """
    Processes all .txt files in a folder.
    Returns a list of results.
    """
    results = []
    folder = Path(folder_path)
    
    for file_path in folder.glob("*.txt"):
        print(f"Processing: {file_path.name}")
        
        text = file_path.read_text(encoding="utf-8")
        
        # Step 1: Classification (a fast small model)
        doc_type = classify_document(text, model="llama3.2")
        print(f"  Type: {doc_type}")
        
        # Step 2: Information extraction (a better model)
        info = extract_info_from_document(text, model="mistral:7b")
        
        results.append({
            "file": file_path.name,
            "classified_type": doc_type,
            "extracted_info": info
        })
        
        print(f"  Done: {json.dumps(info, ensure_ascii=False, indent=2)[:200]}...")
    
    return results


# Demo mode if there are no real documents
def demo():
    sample_text = """
    LEASE AGREEMENT No. 123
    
    Cuenca, January 15, 2026
    
    Ecuador Realty LLC (Landlord) and John Smith (Tenant)
    have entered into this agreement as follows:
    
    1. The Landlord leases an apartment of 700 sq ft at: 45 Example St.
    2. Monthly payment: $500 USD.
    3. Lease term: 12 months starting February 1, 2026.
    4. The Tenant agrees to pay rent on time.
    """
    
    print("=== DEMO: Processing a document with a local model ===\n")
    print("Document:")
    print(sample_text[:200] + "...\n")
    
    print("Classification (llama3.2)...")
    doc_type = classify_document(sample_text)
    print(f"Type: {doc_type}\n")
    
    print("Information extraction (mistral:7b)...")
    info = extract_info_from_document(sample_text)
    print("Result:")
    print(json.dumps(info, ensure_ascii=False, indent=2))
    
    print("\n✅ All data was processed locally and didn't go anywhere!")


if __name__ == "__main__":
    import sys
    
    if len(sys.argv) > 1:
        # If a folder was passed, process it
        folder = sys.argv[1]
        results = process_folder(folder)
        
        output_file = "results.json"
        with open(output_file, "w", encoding="utf-8") as f:
            json.dump(results, f, ensure_ascii=False, indent=2)
        
        print(f"\n✅ Results saved to {output_file}")
    else:
        # Otherwise, the demo
        demo()

Running it:

bash
# Make sure Ollama is running and the models are downloaded
ollama list  # llama3.2 and mistral:7b should be there

# Demo with a test document
python document_processor.py

# Process a folder of documents
python document_processor.py /path/to/documents/

Step 3: Add it to Claude Code through CLAUDE.md

Add this to your project's CLAUDE.md:

Type this into the chat
## Local AI stack

When working with confidential client data, use Ollama (local).
API endpoint: http://localhost:11434/v1
Available models: mistral:7b (general), qwen2.5-coder (code), llama3.2 (fast)

Call pattern: openai.OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

Common mistakes

Mistake 1: Ollama isn't responding

bash
# Check that the service is running
curl http://localhost:11434/api/tags

# If not, start it
ollama serve

# On macOS, Ollama runs as an app (an icon in the menu bar)
# or through launchctl

Mistake 2: The model responds very slowly

If it's under 3-5 tokens/sec, the model doesn't fit in RAM and is spilling into swap. The fix: pick a smaller model or a lower quantization.

bash
# Check how much RAM Ollama is using
ollama ps  # Shows active models and memory

# Pick a smaller model
ollama run llama3.2  # 3B instead of 7B

Mistake 3: The model's JSON won't parse

Local models sometimes add a markdown wrapper. Always handle it:

python
def clean_json(text: str) -> str:
    """Strips the markdown wrapper from the model's response"""
    text = text.strip()
    if "```json" in text:
        text = text.split("```json")[1].split("```")[0]
    elif "```" in text:
        text = text.split("```")[1].split("```")[0]
    return text.strip()

Mistake 4: Different results from the same prompt

That's normal at a high temperature. For structured output (JSON, classification), use temperature=0.0 or temperature=0.1.


Lesson summary

What you can do now:

  • Install Ollama and run a local model in a matter of minutes
  • Pick the right model for the task (size, specialization)
  • Integrate a local model into Python code through the OpenAI-compatible API
  • Implement the LLM Router pattern: different tasks, different models
  • Process a folder of confidential documents locally, without sending data to the cloud

The main idea:

Local models aren't a replacement for Claude. They're an extra tool in your arsenal. A smart entrepreneur uses both: local where you need privacy or volume, the cloud where you need quality.

🎨 Final picture: an LLM Router is like a smart project manager. They know: a simple task → give it to a junior (fast and cheap), a complex one → give it to a senior (high quality but pricier), a confidential one → only to an in-house employee (no outsiders). That's exactly how a well-built AI stack works.


Homework

Level 1 (required):

  1. Install Ollama
  2. Download mistral:7b and llama3.2
  3. Run the demo script from the lesson
  4. Make sure processing works locally (it should work without internet)

Level 2 (recommended):

  1. Take a real document from your business (with no confidential data, for the test)
  2. Adapt the prompt in extract_info_from_document to your document type
  3. Process 5-10 documents and check the quality
  4. Compare the result with the same prompt through Claude Sonnet, and write down the difference in quality

Level 3 (advanced):

  1. Implement smart_router.py from the theory section
  2. Add logging: model used / response time / approximate cost (for the cloud)
  3. Find a task in your business where a local model is good enough, and calculate the monthly savings


Next lesson: Fine-tuning: when prompts aren't enough

The mark stays in this browser only and is never sent anywhere. My progress