Imagine you've hired a smart, versatile employee. For most tasks, giving them instructions is enough, and they'll handle it. For complex tasks, you hand them a folder of documents: let them read and answer. But sometimes a task is so specialized that the only way forward is to send the person to a three-month retraining program. Fine-tuning (further training a model on specialized data) is that "retraining." It's expensive and slow, but afterward the person works on autopilot in that area. In this lesson we'll look at when this path is really needed and when it's a waste of money.
The gist
Most developers jump to fine-tuning when the AI "doesn't get" a task. That's a mistake. In most cases the problem is solved by a better prompt or by adding the right documents. Fine-tuning is a last-resort tool: powerful, expensive, and needed only in specific situations.
This lesson gives you a clear decision tree: what to choose and when.
Key concepts
- Three levels of customizing AI: from simple to expensive
- Prompt engineering (the craft of writing effective requests to AI): when it's enough
- RAG (Retrieval-Augmented Generation, adding outside data to a request): when you need documents
- Fine-tuning: when you need a specialized model
- A decision tree: what to choose for your task
- The cost of each level: orders of magnitude
Theory
Three levels of customizing AI
Level 1: Give instructions. You say: "Write posts in this tone, here are examples, follow these rules." Most employees will manage. Free, instant.
Level 2: Give them a reference library. The employee doesn't know all the company's documentation, but you give them access to the archive: "Before answering a customer, find the right section and use it." It takes setup and costs money, but the data lives separately.
Level 3: Send them to a specialized training program. Three months of training on nothing but your specifics. Afterward the person knows everything by heart and works fast. Expensive and slow, but sometimes the only way.
Prompt engineering, RAG and fine-tuning are exactly these three levels.
Level 1: Prompt engineering, when it's enough
What it is: writing detailed system prompts, examples and a response structure right in the request to the model.
When it's enough:
- Most business tasks
- You need a particular writing style
- Document analysis, translation, code from a template
- Classification, summarizing, question-answering
Advantages:
- Cost: just the time to write the prompt (plus the usual API price)
- You can change it instantly: edit the text and you're done
- No need to retrain the model
- Flexible: different tasks, different prompts
Limitations:
- When the task requires knowledge the base model doesn't have
- When the style is so specific it can't be described in words
- When there are thousands of identical tasks and the prompt tokens add up to real money
Examples of tasks where a prompt is enough:
- Writing posts in a specific brand's style (describe the style + give 3 examples)
- Reviewing a contract against a checklist (the checklist goes in the prompt)
- Translation that respects terminology (the terms go in the prompt)
- Generating code from a template (the template goes in the prompt)
The rule: always start with a prompt. Add complexity only if the prompt can't handle it.
Level 2: RAG, when the knowledge is out of date or too large
What it is: before answering, the AI searches your knowledge base for relevant chunks of text and adds them to the request's context.
How it works technically (no math):
The user's question
↓
Search in a vectorDB (a vector database: a store of texts
as numeric vectors that you search for similar ones)
↓
The top 3 relevant documents are found
↓
Claude gets: the question + the chunks it found, and answers
↓
An answer with links to the sourcesWhen you need RAG:
- The data changes often (prices, laws, internal company policies)
- A large document base that can't be crammed into a single prompt
- Specialized industry knowledge the base model doesn't have
- You need to cite specific sources
- Confidential data you don't want to duplicate in the prompt
Tools for RAG:
- Gemini Notebook (formerly NotebookLM, renamed in July 2026) from Google: there's a free version; you upload documents and ask questions
- LlamaIndex: a Python library for building RAG systems
- LangChain: a popular framework (a ready-made scaffold for development) for LLM apps
- Anthropic Claude with document uploads: a built-in feature of Claude.ai
The cost of RAG: depends on the tool and the amount of data: from zero with ready-made services to ongoing costs for a vectorDB and compute (computing power)
Examples where RAG works great:
- A company bot that answers from internal documents and policies
- A legal database of court decisions: "find precedents on this issue"
- Medical protocols: "what's the standard of care for this diagnosis?"
- Product customer support: a base of 500 FAQ (frequently asked questions) articles
Level 3: Fine-tuning, when you really need specialization
What it is: retraining a model on a specialized dataset. The model doesn't just "remember" the examples; it physically changes the neural network's weights, meaning the model's internal math changes.
When fine-tuning is really justified:
A very specific style that can't be described in a prompt. For example, a particular author's voice, with thousands of nuances in word choice.
A huge number of identical tasks. If you have 100,000 requests a day of one type, every long prompt full of instructions costs extra tokens. After fine-tuning the prompt is shorter and inference (the process of the model generating an answer, and its speed) is faster.
Confidential data. Train it on local hardware, and the data doesn't go anywhere.
Specialized terminology or language. Medical abbreviations, legal phrasing, the professional jargon of a narrow industry.
You need fast inference with a small model. A small trained model can run faster and cheaper than a large general-purpose one.
When you DON'T need fine-tuning (common mistakes):
| What the developer thinks | What's actually needed |
|---|---|
| "I want the AI to know our documents" | RAG: that's its job |
| "The model gives wrong answers" | Improve the prompt |
| "I want to keep the conversation context" | Memory / context management |
| "I need a specific style" | 3-5 examples in the prompt |
| "The model is slow" | Pick a faster model |
How fine-tuning works technically (no math)
Step 1: The base model. You take a ready, already-trained model. Today that's usually an open model (the Llama, Qwen, Gemma or Mistral families), depending on the task and budget. Closed-model vendors are phasing this option out.
Step 2: You prepare a dataset (the data set for training). "Question → correct answer" pairs, or "text → the right classification." For a basic result you need dozens of examples; for a good one, hundreds and thousands. The exact number depends on the task.
Step 3: You start the training. The model looks at your examples, compares its answers with the correct ones and gradually adjusts its internal parameters.
Step 4: You get a new version of the model. The same architecture, but "tilted" toward your data. It does better on your task and worse on everything else (that's normal).
The main risk is overfitting (when the model memorizes the examples but doesn't learn to generalize):
How to avoid it: split the dataset into train/test sets, and watch the metrics (numeric quality measures) on both parts.
PEFT and LoRA: efficient fine-tuning without huge costs
The problem with full fine-tuning: a large language model has billions of parameters (the numbers inside a neural network that determine its behavior). Changing all of them takes enormous computing resources.
The solution is PEFT (Parameter-Efficient Fine-Tuning): you change only a small fraction of the parameters instead of all of them. The quality is often nearly the same, and the cost is much lower.
The most popular PEFT method is LoRA (Low-Rank Adaptation):
Instead of changing the whole huge matrix of parameters, you add a small "layer" on top. That layer gets trained on your data. The base model isn't touched.
The result:
- Quality usually close to full fine-tuning
- Many times cheaper
- You can switch between different LoRA adapters (add-on layers for different tasks) on one base model
Most modern open-source fine-tuning services use LoRA under the hood.
Where to fine-tune: comparing providers
Important as of October 2026: the big closed-model vendors are winding down self-service fine-tuning. OpenAI notified developers about this in May 2026: organizations that hadn't fine-tuned models before can't create training jobs as of May 7, 2026, and existing customers lose the ability to create new jobs on January 6, 2027. Fine-tuned gpt-3.5-turbo and gpt-4 models are being shut down on October 23, 2026. So for a new project, open models that you can fine-tune yourself or through a GPU provider are the safer bet. Providers' terms change, so check them on their websites.
| Service | Models | Terms | Difficulty | Best for |
|---|---|---|---|---|
| OpenAI Fine-tuning | OpenAI models | Self-service fine-tuning is being phased out (dates above) | — | Don't choose it for a new project |
| Anthropic | Claude | There's no self-service fine-tuning of Claude's weights; discuss custom solutions for large businesses with Anthropic directly | Enterprise | Large businesses |
| Google Cloud (Vertex AI) | Gemini and others | See Google's documentation for the list of models and terms | High | Google infrastructure |
| Hugging Face | Llama, Mistral, Qwen, Gemma and more | You pay for GPU time; prices depend on the hardware | High | Privacy, open source |
| Together AI | Open-source models | Check the current terms on the website | Medium | Price/quality balance |
| Replicate | Various | Check the current terms on the website | Low | Quick experiments |
Recommendations:
- Beginners: first make sure a prompt and RAG really aren't enough. If you do need fine-tuning, use an open model + LoRA through Hugging Face (it has no-code AutoTrain) or another service where fine-tuning is currently available. Work through a tutorial example before you spend money.
- Privacy matters: Hugging Face + your own hardware or a rented GPU (graphics processing unit, used to train neural networks), so the data doesn't go anywhere
- Limited budget: an open model (Llama, Qwen, Gemma, Mistral) + LoRA on a rented GPU gives good value for money
- Enterprise: discuss custom solutions directly with the model vendor, if the budget allows and you truly can't do without it
The decision tree: what to choose?
I want the AI to handle my task better
│
├── Does the data change often, or is there a very large amount?
│ └── YES → RAG (a vector database)
│ (prices, laws, company documents, FAQs)
│
├── Is the data static and small (fits in a prompt)?
│ └── Try prompt engineering first
│ │
│ └── The prompt isn't working?
│ │
│ ├── Is the style too specific to describe?
│ │ └── YES → Fine-tuning
│ │
│ ├── Specialized terminology or language?
│ │ └── YES → Fine-tuning
│ │
│ ├── 100K+ requests a day of one type?
│ │ └── YES → Fine-tuning (saves on tokens)
│ │
│ └── NONE of the above →
│ Improve the prompt, add more examples
│
├── Is the data confidential (can't go to the cloud)?
│ └── YES → Fine-tuning on your own hardware
│ (Hugging Face + Llama + LoRA)
│
└── Need maximum speed / minimum price per request?
└── YES → Fine-tune a small model
(a small trained model is faster than a large general one)A simplified rule for 90% of cases:
Tried a prompt → it doesn't work
↓
Need outside data? → RAG
↓
Still doesn't work → Fine-tuningCost: orders of magnitude
Specific amounts go out of date quickly and depend on the provider, so this section describes orders of magnitude. For current API prices, see the What's current page and the providers' websites.
Prompt engineering:
- Cost: the time to write a good prompt (hours)
- Operational cost: the standard API price
RAG:
- Setup: from zero (ready-made services like Gemini Notebook) to noticeable development costs
- Ongoing: vectorDB, compute, API price
- Time to results: days
- Operational cost: the standard API price + vectorDB
Fine-tuning (open source, Hugging Face + GPU):
- GPU rental: by the hour; the price depends on the graphics card
- One training run: noticeably more expensive than a prompt or RAG, and you usually need several iterations
- Inference: cheap (your own infrastructure) or nearly free (your own hardware)
- Time to results: several days or more (data preparation + training)
Fine-tuning with closed-model vendors: the terms change quickly (see the table above); for a large business, get the price and timeline from the vendor.
The takeaway for most people: start with prompts (cheapest) → add RAG (ongoing costs) → fine-tuning only when there's a real business need (most expensive and slowest). The lessons What a client costs and what they bring in and Break-even and your cash cushion show how to figure out whether it will pay off.
When fine-tuning is really justified: examples
A call center with 10,000 calls a day:
- Task: automatically classify the topics of incoming calls
- Why fine-tuning: a huge volume → a long prompt every time is expensive; speed matters; the dataset (call transcripts, the text versions of the audio) already exists
- Result: a small model works fast, cheaply and accurately for this specific task
Medical documentation:
- Task: structure doctors' discharge summaries and fill in fields in a CRM (customer relationship management system)
- Why fine-tuning: specialized medical abbreviations the base model doesn't have; patient confidentiality; high accuracy is critical
- Result: a model trained on thousands of real discharge summaries is significantly more accurate
A law firm:
- Task: analyze contracts, find risky wording
- Why fine-tuning: specialized legal constructs of a particular legal system (for example, US state law or Latin American civil law); nuances you can't describe in a prompt; confidentiality
- Result: the model "thinks" like an experienced lawyer in that jurisdiction (legal territory)
A game studio:
- Task: dialogue for NPCs (non-player characters, the computer-controlled characters) with each character's unique voice
- Why fine-tuning: thousands of lines of dialogue, and each NPC has its own style that can't be described in a prompt
- Result: NPCs speak "in character" without a big prompt every time
Examples where fine-tuning ISN'T needed (but people thought it was):
- "I want the bot to know about our company" → load the documents into RAG; Gemini Notebook will handle it quickly
- "Claude doesn't understand our terminology" → add a glossary (a list of terms) to the system prompt
- "I want it to write like our brand" → 5-10 sample texts in the prompt + a description of the style
If you've decided to fine-tune: a basic checklist
Step 1: Prepare the dataset
- OpenAI-style conversation format: JSONL (a file format where each line is a JSON object), which many tools for open models understand too
- Dozens of examples to start; hundreds or more for a good result
- Quality matters more than quantity: garbage in, garbage out
- Split it into train/validation sets (80/20)
An example structure (conversation format):
{"messages": [
{"role": "system", "content": "You are a support agent for company X"},
{"role": "user", "content": "How do I get a refund?"},
{"role": "assistant", "content": "To request a refund, email [email protected]..."}
]}Step 2: Choose the model and provider
- Beginners: a small open model + LoRA on a service with a ready-made interface
- Privacy: Hugging Face + an open model + LoRA on your own hardware
Step 3: Start the training
- Hugging Face: through AutoTrain, a no-code interface for fine-tuning
- Your own GPU or a rented one: libraries like Unsloth for LoRA (see the Local AI models lesson)
Step 4: Evaluate the result
- Compare it with the base model on a test dataset
- Test edge cases (unusual inputs)
- Check for overfitting: does it work well on new data?
Step 5: Iterate (repeat the improvement cycle)
- Add more examples if the quality isn't good enough
- Adjust the hyperparameters (the settings of the training process)
- Sometimes improving the dataset is cheaper than training longer
Practice
Task 1: Figure out what your task needs
Take a real task you're currently solving with Claude. Go through the decision tree:
- Write the task in one sentence
- Answer the questions in the decision tree above
- Decide: prompt engineering / RAG / fine-tuning
- Estimate the budget and time
Task 2: Try NotebookLM (free RAG)
If you have a knowledge base (PDFs, Word documents, web pages):
- Open Gemini Notebook (at notebook.google, formerly NotebookLM)
- Create a new notebook
- Upload 5-10 documents from your company or project
- Ask questions and see how RAG finds answers in the documents
- Compare the quality of the answers with plain Claude without the documents
Result: in 20 minutes you'll have a working RAG setup without a single line of code. This helps you see whether you need fine-tuning or whether RAG already solves the problem.
Task 3: Work through a LoRA tutorial
If you want to try fine-tuning:
- Open Hugging Face's learning materials on LoRA and PEFT (huggingface.co/learn and the peft library documentation)
- Follow the example with sentiment classification (deciding whether a text is positive, negative or neutral) on a small open model
- You need: a free notebook like Google Colab, or a GPU rented for a couple of hours (see the websites for prices)
- Time: a few hours, including data preparation
Note: this task used to use OpenAI fine-tuning, but OpenAI is phasing out self-service fine-tuning (see the provider table above).
The goal of this task isn't the result but understanding the process. After it, your decision about "do I need fine-tuning or not" will be an informed one.
Key takeaways
- Most tasks are solved with a prompt. Always start there. It's cheap and fast.
- RAG is the second step, when you need outside documents, frequently updated data or large knowledge bases. You can try Gemini Notebook for free.
- Fine-tuning only when you really need it: a specific style, confidentiality, huge volumes of identical tasks, specialized language/terminology.
- LoRA makes fine-tuning more accessible: quality close to full training at a much lower cost.
- For most practitioners: prompt engineering + RAG is all you'll need for years to come.
- The decision tree: data changes often → RAG. Static and can't be described in a prompt → fine-tuning. Otherwise → improve the prompt.
Next lesson
Local AI Agents 2026: nanoClaude, OpenClaw, Hermes, Qwen-Agent
AI Mayak Academy | Advanced techniques
The mark stays in this browser only and is never sent anywhere. My progress