Library · Memory and context: keeping the agent on track

RAG: Retrieval Augmented Generation

Builder60 minUpdated: October 2026
21 of 105 in the library

Module: Tokens and context | Time: about 30 min theory + 30 min practice


The gist

Claude is smart, but it knows nothing about your company, your clients or your documents. RAG is like giving Claude access to your personal library: it can find the right book in a second and use it when answering. Without RAG, it answers from general knowledge. With RAG, it answers from your specific data.


Key concepts

  • RAG solves the "the model doesn't know my data" problem
  • Data gets turned into vectors: mathematical fingerprints of meaning
  • Search works by meaning (semantic), not by keywords
  • Stack: Gemini Embedding 2 (vectorization) + Pinecone (storage) + Claude (generation)

Theory

The problem RAG solves

🎨 Picture this: Claude without RAG is like a walking encyclopedia invited to consult for your company. It knows everything about the world but nothing about your clients, your prices, your rules. RAG is the folder of documents you hand over before the meeting.

Claude was trained on data up to a certain date and knows nothing about your business. If you build a consultant agent for a client, the agent doesn't know the company's products, internal rules or client history. There are several ways to solve this:

Option 1: Cram everything into the system prompt Put all the documentation right into claude.md. The problems: the context window limit (as of October 2026, Opus 5.5 and Sonnet 5.5 have 1 million tokens, Haiku 4.5 has a smaller window, see What's current; a large corporate knowledge base still won't fit, and answer quality drops long before the window is full), the cost (it's read on every request), and it's hard to keep updated. For a small knowledge base, on the other hand, this is the simplest path: if you don't have much material, put it into the context and don't build extra infrastructure.

Option 2: Fine-tuning Train the model further on your data. Expensive, slow, requires expertise, and the result is unpredictable. Not for most tasks.

Option 3: RAG Store the data separately, search for what's relevant on each request, and add what you found to the context. This is the right approach for most tasks involving internal data.


How RAG works: step by step

Step 1: Preparing the data (Indexing)

🎨 Picture this: an embedding is like building a scent map for a bloodhound. Text gets turned into a "digital scent," a vector. Texts with similar meaning smell similar. The dog (the search system) finds what it needs by scent, not by words.

You take your data (PDF documents, articles, web pages, spreadsheets, images, videos) and run it through an embedding model.

The embedding model turns each piece of text (or image, or video) into a vector: an array of thousands of numbers. These numbers encode the meaning of the content. Texts with similar meaning get similar vectors, even if they're written in different words.

Example: "How do I cancel my subscription" and "Procedure for terminating the agreement" use different words but mean the same thing. Their vectors will be close.

Step 2: Storing the vectors (Vector Database)

🎨 Picture this: a vector database is a library where books are shelved by meaning, not alphabetically. Books about "love" sit next to each other even if they're written in different languages. A regular database can't do that.

All the resulting vectors are saved in a vector database, Pinecone. It's a specialized database optimized specifically for searching vectors.

Each vector is stored along with the original text and metadata (source, date, category). That way, once a vector is found, you can pull up the original text.

Step 3: Search (Retrieval)

When a user asks a question, the question is also turned into a vector (by the same embedding model). Then the database looks for the N closest vectors by mathematical distance (cosine similarity).

The result: the top 3 or top 5 documents that are semantically closest to the question.

Step 4: Augment + Generate

The documents that were found are added to the context before the request to Claude:

Code
Here are relevant documents from the knowledge base:
[Document 1: Product return policy...]
[Document 2: FAQ on canceling a subscription...]

Answer the user's question using the information from these documents:
"How do I return an item I bought 3 months ago?"

Claude sees the specific data and answers based on it, not from general knowledge.


Google Gemini Embedding 2: multimodality

🎨 Picture this: multimodal embeddings are like an interpreter who understands not only text but also gestures, paintings and music. Show them a photo of a broken part, and they'll find a similar one in the catalog without a single word.

For vectorization, this stack uses Google Gemini Embedding 2 (model ID gemini-embedding-2), an embedding model with multimodal capabilities. As of October 2026 it has moved from preview to stable. There's also the previous, text-only model gemini-embedding-001.

What this means in practice:

  • Text documents → vector ✅
  • Images (JPG, PNG) → vector ✅
  • Video → vector ✅
  • Audio → vector ✅
  • PDF documents → vector ✅

This opens up possibilities that text-only embedding models don't have: a knowledge base with product images, video tutorials, voice notes, all of it becomes searchable.

Example: an appliance manufacturer stores photos of every model. A user uploads a photo of a broken device, and the system finds similar images in the database and identifies the model even without a name.


Pinecone: why a vector database and not a regular one

A regular database (PostgreSQL, MySQL) can search for exact matches: "find records where category = 'FAQ'". It can't search by meaning.

Pinecone is optimized for one task: "find the N vectors closest to this vector." Search takes milliseconds even across millions of vectors. It's a cloud-based solution, so you don't have to set up your own infrastructure.

Alternatives: Weaviate (open source), Qdrant (open source, can be self-hosted), pgvector (an extension for PostgreSQL), Chroma (local, for development). Pinecone is good for getting started: simple integration, and there's a free Starter plan with limits on volume and number of operations.


How to build a RAG system

  1. Prepare the data: gather the documents, articles, FAQs and images the system needs to know
  2. Split it into chunks: large documents get broken into pieces of about 500 words (with an overlap of about 50 words). This matters: a whole document in one vector works worse than several chunks on specific topics
  3. Create an index in Pinecone: through the API you create a "space" for storage
  4. Upload the vectors: for each chunk you get a vector from Gemini Embedding 2 and upload it to Pinecone
  5. Build a query interface: a function that takes a question, searches Pinecone and adds the results to Claude's context
  6. Test: ask questions, check the quality of the answers, adjust

When you need RAG and when you don't

You need RAG for:

  • A consultant agent for a company's products/services
  • Searching a corporate knowledge base
  • Questions about specific documents (legal, technical)
  • Personalized recommendations based on history

You don't need RAG for:

  • Tasks where Claude already knows everything (general questions, programming)
  • When all the information you need fits into claude.md
  • One-off tasks where the data doesn't change

Real-world uses of RAG

Company knowledge base: an employee asks "how do I file travel expenses?" The agent finds the right section of the HR policy and explains it in its own words.

Online store consultant agent: a shopper describes a problem with a product. The agent finds similar cases in the solutions database and suggests a specific next step.

Legal assistant: a lawyer uploads a batch of contracts and asks "do these contracts have force majeure clauses?" The system finds the relevant sections.

Ticket-based support: the agent has seen thousands of similar questions from customers. It finds comparable ones and uses the answers that already helped.


Practice

Assignment: a simple RAG system over text

For practice, we'll build a small RAG system with no external services, using local storage.

  1. Create a knowledge-base/ folder in your project with 5-7 text files on a topic you know well (for example, an FAQ about your services)
  2. Ask Claude Code: "Build a simple RAG system that searches the files in knowledge-base/. On each query, it finds the most relevant document and answers based on it. Use TF-IDF for search (no external APIs)."
  3. Test it: ask a question that's covered in one of the documents, then a question that isn't
  4. For the full stack with Pinecone + Gemini Embedding: create an account at pinecone.io (free Starter plan), get an API key (keep it in .env, not in the chat), and ask the agent to migrate the system

Comparing embedding models

Model Modalities Cost (as of October 2026) Best for
Google Gemini Embedding 2 Text + images + video + audio + PDF There's a free tier with limits; paid prices are on the Gemini API pricing page Multimodal knowledge bases
Voyage AI (voyage-4 family) Text (high quality), multimodal models available Has a free allowance; current prices on the Voyage AI site Precise semantic search over text
OpenAI text-embedding-3-small Text Low per-token price; see OpenAI's pricing page Budget option, text
OpenAI text-embedding-3-large Text (high quality) Higher than the small model; see OpenAI's pricing page Maximum accuracy on text

Comparing vector databases

Free plan terms change: check current prices and versions on the services' websites and on the What's current page.

Database Type Free plan Best for
Pinecone Cloud Starter plan with limits Production, simple integration
Chroma Local Free Development, prototypes
Qdrant Self-host / Cloud Free (self-host) Full control, open source
Weaviate Self-host / Cloud Self-host is free; for cloud see the pricing page Complex queries, GraphQL
pgvector PostgreSQL extension Free If you already have PostgreSQL

Tools and resources

  • Pinecone: a vector database with a free starter plan
  • Google Gemini Embedding 2: via the Google AI API, a multimodal embedding model
  • Anthropic: Embeddings: the official guide to embeddings for Claude (with a link to a RAG example using Pinecone)
  • LangChain / LlamaIndex: Python frameworks that make it easier to build RAG pipelines
  • Chroma: a local vector database for development (no sign-up)
  • Voyage AI: a high-accuracy embedding model; Anthropic doesn't have its own embedding model, and Anthropic's documentation points to Voyage
  • OpenAI text-embedding-3-small: an alternative embedding model (text only, cheaper)

Common mistakes

🎨 Picture this: RAG without testing is like opening a restaurant and deciding everything's fine because the first two customers didn't complain. The tenth customer will order something that's not on the menu, and that's when the trouble starts.

Mistake 1: Chunks that are too big You uploaded whole 5,000-word documents as one vector each. The result: search finds the document but the wrong part of it. The best chunk size: 300-500 words with an overlap of 50-100 words.

Mistake 2: RAG when CLAUDE.md is enough If you have 10 rules and 5 templates, put them in CLAUDE.md. You need RAG when you have noticeably more data than it's reasonable to keep in context (as of October 2026 the window is 1M tokens for Opus 5.5 and Sonnet 5.5, smaller for Haiku 4.5, and quality drops before the window fills up). For small knowledge bases, RAG is overkill.

Mistake 3: Not testing answer quality You set up RAG, checked one question, and it works. But out of 10 different questions, 3 answers are inaccurate. Create a set of 15-20 test questions with correct answers and check quality systematically.


Cross-references


Key takeaways

RAG is the bridge between Claude's vast knowledge and your business's specific data. Without RAG, the agent is smart but blind to your particular reality.

Vector search looks for meaning, not words. That's what fundamentally sets RAG apart from a simple grep or Ctrl+F: users can ask in their own words and still find what they need.

Multimodality (text + images + video) opens up uses that aren't possible with text embeddings. It's a new capability, and it's worth trying on your own materials.


Next lesson

→ Websites and web apps from scratch: from a prompt to a finished website

The mark stays in this browser only and is never sent anywhere. My progress