What RAG is, and when you need it
Before building anything, be clear about which problem of a language model RAG solves — and which problems it does not.
Three limits of an LLM
A language model learns from data up to a cut-off date and is then frozen. Put one in front of real users and three limits show up almost immediately:
- It has never seen your data. Internal policies, contracts, product manuals, support tickets — none of it was in the training set.
- Its knowledge is stale. A price list that changed last week, a policy issued yesterday: the weights cannot know.
- It invents, and it cites nothing. When it is unsure the model still answers fluently, and the reader has no way to check the claim.
The idea behind RAG
Retrieval-Augmented Generation (Lewis et al., 2020) takes knowing out of the model. Instead of hoping the weights remember everything, the system looks up the passages relevant to the question, puts them in the prompt, and asks the model to answer using only those passages.
What you get in return:
- Answers about private and current data — update the corpus, no retraining.
- Answers with citations, so a reader can verify them.
- Access control that actually works: each user retrieves only the documents they are allowed to see.
Two phases: indexing and querying
Indexing — offline
Querying — once per question
The final quality depends on both phases. A pattern you will meet often: most wrong answers do not come from the generator at all, but from the right passage never being retrieved — which is why chunking, embedding and retrieval get the most space in this course.
One example, end to end
Question: Can an employee on probation work remotely?
[Retrieval] top-3 passages from "Remote work policy v2.1":
[1] "Employees who have completed their probation may register for up to 2 days/week..."
[2] "During probation, remote work requires case-by-case approval from the line manager..."
[3] "Requests are submitted through the HR portal before 17:00 on Friday..."
[Generation] An employee on probation cannot register a fixed schedule, but may work
remotely with case-by-case approval from their line manager [2]. After
probation, up to 2 days per week [1].
RAG, fine-tuning, or a long context?
| RAG | Fine-tuning | Whole corpus in the context | |
|---|---|---|---|
| Adding new knowledge | Good — update the corpus and you are done | Poor; expensive, and it forgets | Good if the corpus is small |
| Teaching style or format | Limited | Good | Limited |
| Citing sources | Natural | No | Possible, but hard to pin down |
| Large corpus (GB) | Yes | — | Does not fit, and is expensive |
| Per-user permissions | Filter at retrieval time | No | You must filter beforehand |
| Cost per question | Low to moderate | Low | High once the context is long |