What RAG is, and when you need it

Before building anything, be clear about which problem of a language model RAG solves — and which problems it does not.

Three limits of an LLM

A language model learns from data up to a cut-off date and is then frozen. Put one in front of real users and three limits show up almost immediately:

The idea behind RAG

Retrieval-Augmented Generation (Lewis et al., 2020) takes knowing out of the model. Instead of hoping the weights remember everything, the system looks up the passages relevant to the question, puts them in the prompt, and asks the model to answer using only those passages.

A way to picture it A bare LLM is a candidate in an oral exam working from memory. An LLM with RAG is the same candidate allowed to bring the documents in: understanding and phrasing are still on them, but the facts come off the page and they can point to which page.

What you get in return:

Two phases: indexing and querying

Indexing — offline

LoadPDF, HTML, wiki, DB
→
Cleanstrip headers, footers, noise
→
Chunksplit into passages
→
Embedpassage → vector
→
Storevectors + metadata

Querying — once per question

Ask“How much annual leave do I get?”
→
Retrievetop-k nearest passages
→
Rerankkeep the best ones
→
Promptinstructions + sources + question
→
Generateanswer with [1], [2]

The final quality depends on both phases. A pattern you will meet often: most wrong answers do not come from the generator at all, but from the right passage never being retrieved — which is why chunking, embedding and retrieval get the most space in this course.

One example, end to end

Question: Can an employee on probation work remotely?

[Retrieval] top-3 passages from "Remote work policy v2.1":
  [1] "Employees who have completed their probation may register for up to 2 days/week..."
  [2] "During probation, remote work requires case-by-case approval from the line manager..."
  [3] "Requests are submitted through the HR portal before 17:00 on Friday..."

[Generation] An employee on probation cannot register a fixed schedule, but may work
             remotely with case-by-case approval from their line manager [2]. After
             probation, up to 2 days per week [1].

RAG, fine-tuning, or a long context?

RAGFine-tuningWhole corpus in the context
Adding new knowledgeGood — update the corpus and you are donePoor; expensive, and it forgetsGood if the corpus is small
Teaching style or formatLimitedGoodLimited
Citing sourcesNaturalNoPossible, but hard to pin down
Large corpus (GB)Yes—Does not fit, and is expensive
Per-user permissionsFilter at retrieval timeNoYou must filter beforehand
Cost per questionLow to moderateLowHigh once the context is long
A quick rule Need knowledge (facts, documents, numbers that change) → RAG. Need behaviour (tone of voice, output format, a reasoning procedure of your own) → fine-tuning. A small corpus that rarely changes, with few users → try putting it straight in the context before you build a whole pipeline.
RAG is not a cure-all A question that requires reading the entire corpus (“what are this year's complaint trends?”) or arithmetic over structured data (“Q3 revenue?”) usually needs a SQL query, an analytics step, or one of the variants in Common failures & what is next — not top-k passages.

Check yourself

Why does RAG reduce hallucination without eliminating it? Because the model can still ignore the context, over-infer from it, or be handed context that simply does not contain the answer. You need a prompt that forces grounding, permission to say “I don't know”, and a faithfulness measurement (see Building the prompt and Measuring quality).
A company wants a chatbot that speaks in its brand voice and answers from 5,000 pages updated weekly. What do they use? Both, in different roles: RAG for the knowledge in 5,000 weekly-changing pages; the brand voice is usually reachable through the system prompt, and fine-tuning is worth it only once the prompt demonstrably is not enough.
An answer is wrong. What do you inspect first? The retrieved passages: is the answer in them at all? If it is not, the fault is in retrieval or indexing; if it is there and the answer is still wrong, the fault is in the prompt or the generator.