Two different problems
A language model can fall short in two ways. It can lack knowledge: it does not know your products, your policies, your contracts or what changed last week. Or it can behave wrongly: it knows enough, but answers in the wrong format, the wrong tone, or makes the same kind of mistake on a task again and again.
Retrieval-augmented generation, usually called RAG, fixes the first problem. Fine-tuning fixes the second. Using one to solve the other is the most common and most expensive mistake in AI projects.
The confusion is understandable, because both are described as giving the model your data. The difference is where the data ends up. With retrieval it stays in your documents and is fetched fresh for each question. With fine-tuning it is absorbed into the model's weights, where it shapes how the model writes but cannot be pointed to, corrected or removed without training again.
| Symptom | Likely problem | Likely fix |
|---|---|---|
| It does not know our products, prices or policies | Knowledge | Retrieval |
| It gives answers that were true last year | Knowledge | Retrieval |
| We need to see where each answer came from | Knowledge | Retrieval |
| It ignores our report format however we ask | Behaviour | Better prompts, then fine-tuning |
| It labels tickets inconsistently | Behaviour | Fine-tuning on labelled examples |
| It is right but sounds nothing like us | Behaviour | Fine-tuning on examples of our writing |
What retrieval does
With retrieval, the model does not memorise your documents. When a question arrives, the system searches your documents for the passages most likely to answer it and hands them to the model with the question. The model answers from what it was given, and the application shows which passages it used.
The approach was set out in a 2020 research paper which noted that, for models relying only on what they learned in training, providing provenance for their answers and updating their knowledge were open problems, and which paired a model with an explicit, searchable memory to address them [1]. Those are exactly the two properties businesses care about: answers you can check, and knowledge that is current the moment a document changes.
- 01
Collect and clean
Gather the documents the model should answer from, and decide who may see which. Out-of-date and duplicate documents are the main cause of bad answers.
- 02
Split and index
Break documents into passages and store a numerical representation of each for search. pgvector, for example, adds vector similarity search to Postgres, so the index can live in a database you already run [4].
- 03
Retrieve
For each question, find the best passages, usually by combining meaning-based and keyword search, and respect the asker's permissions.
- 04
Answer with sources
Give the passages to the model with instructions to answer only from them and to cite them, and show the citations to the user.
- 05
Keep in sync
Re-index documents when they change, and remove them when they are retired.
What makes retrieval good or bad
When a retrieval system gives a poor answer, the model is rarely the culprit. Far more often, the right passage was never found, or was found alongside several wrong ones. The work that improves answers is mostly search work.
- Splitting documents sensibly: by section and heading rather than by a fixed length, so a passage carries its own context.
- Keeping metadata: the document's title, date, owner and audience, so search can prefer the current policy over last year's.
- Combining search methods: meaning-based search finds paraphrases, keyword search finds part numbers and names. Most good systems use both.
- Re-ranking: a second, more careful pass over the top results before they reach the model.
- Saying "I don't know": when nothing relevant is found, the assistant should say so and offer a person, rather than answer from general knowledge.
What fine-tuning does
Fine-tuning trains an existing model further on examples of the task: inputs paired with the outputs you want. OpenAI's guide lists uses such as getting a model to format responses consistently, classification, generating content in a specific format, correcting instruction-following failures and matching a tone and style, and it notes that prompting alone may be all you need [2]. That last point matters. Try clear instructions and a few good examples in the prompt before you train anything.
Fine-tuning has become much lighter than it was. Methods such as LoRA train small additional weights instead of the whole model; its authors report reducing the number of trainable parameters by 10,000 times and the GPU memory needed by 3 times, compared with fully fine-tuning a very large model [3]. That makes fine-tuning an open model on your own examples a practical project rather than a research programme.
- It needs examples: a few hundred to a few thousand good input and output pairs, checked by people who know the task.
- It does not reliably teach facts. A fine-tuned model can still invent details, and it cannot tell you where an answer came from.
- It has to be repeated when the task changes, or when you move to a newer base model.
How to choose
- 01
Write the evaluation first
Collect real questions or tasks with good answers. Without them you cannot tell whether any change helped.
- 02
Try prompting
Clear instructions, a defined output format and a few examples solve more behaviour problems than people expect.
- 03
Add retrieval if answers need your knowledge
If the model needs facts it was never trained on, or facts that change, retrieval is the fix. Measure again.
- 04
Fine-tune if behaviour is still wrong
If the model has the right information but keeps getting the format, labels or tone wrong, and you have the examples, fine-tune. Measure again.
- 05
Stop when the evaluation says so
Each step costs effort to build and to maintain. Stop at the first one that meets the bar you set.
When you need both
Some projects need knowledge and behaviour at once. A support assistant may need to answer from the current help centre (retrieval) in the company's voice and a strict structure that the support system can parse (fine-tuning). A contract review tool may need to find the relevant clauses (retrieval) and classify them against a playbook the way the client's own team does (fine-tuning).
In those cases we build retrieval first, because it fixes the larger share of errors and is easier to change, and fine-tune a model to use the retrieved passages the way the task requires. The evaluation set covers both: did it find the right passages, and did it use them correctly.
Living with each choice
The two approaches also differ after launch. A retrieval system is maintained like a search engine: documents are added, changed and retired, and the index follows. Anyone who edits the source documents is, in effect, updating the assistant, which is usually what a business wants. Its running cost is the model plus a modest database.
A fine-tuned model is maintained like a trained colleague: when the task changes, it needs new examples and another training run, and the new version has to pass the evaluation before it replaces the old one. If it is an open model, it also needs GPUs to serve it, every month. Neither is hard, but they are different commitments, and the examples below show the difference in running cost.
Three projects, priced from our rate card
The worked examples below show the range from our rate card today, with the build and the monthly running cost. They are priced when you open the page.
Worked example, priced now
Retrieval on a frontier model
A staff assistant that answers from the company's documents with citations, using a frontier model through its API. Upload, search and permissions included.
- Build
- ≈ US$66,700 to US$102,000, delivered within 18 weeksCAD 95,000 to 145,300
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$558 a monthCAD 795 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$383 a monthCAD 545 a month
Worked example, priced now
A fine-tuned model for one task
An open model fine-tuned on the company's labelled examples to classify and route incoming requests in a fixed format, served on rented GPUs and connected to the existing ticket system.
- Build
- ≈ US$104,000 to US$159,000, delivered within 29 weeksCAD 148,300 to 226,200
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$4,360 a monthCAD 6,210 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$4,190 a monthCAD 5,960 a month
Worked example, priced now
Retrieval and fine-tuning together
A support assistant on a fine-tuned open model that answers from the current help centre with citations, in the company's voice and a structure the support system can read.
- Build
- ≈ US$128,000 to US$195,000, delivered within 33 weeksCAD 182,100 to 277,900
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$4,360 a monthCAD 6,210 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$4,190 a monthCAD 5,960 a month
Compare the monthly lines. The retrieval example pays for tokens as it is used; the two fine-tuned examples pay for a GPU server whether it is busy or not. For a team of a few dozen people asking questions, retrieval on a frontier model is usually the sensible start. Open any example in the estimator to change the traffic or the serving choice and see where the balance shifts for you.
Pitfalls to avoid
- Fine-tuning on your documents and expecting the model to quote them. It will learn the style of your documents, not reliably their contents.
- Indexing everything. A retrieval system over every file on a shared drive finds old, duplicated and contradictory passages. Curate first.
- Ignoring permissions. If a person may not open a document, the assistant must not quote it to them. Filter at retrieval time.
- No evaluation set. Without one, every change is a guess, and a model upgrade can quietly make things worse.
- Forgetting maintenance. Retrieval needs documents kept current; fine-tuning needs retraining when the task or the base model changes.