Measuring RAG quality
“I tried a few questions and it seemed fine” is not an evaluation. This lesson builds a measurable, repeatable process, so that every change to chunking, models or prompts arrives with a number attached.
Two layers to measure
Measure them separately, or you will not know which one broke:
| Layer | The question it answers | Metrics |
|---|---|---|
| Retrieval | Was the passage holding the answer fetched, and fetched high up? | Recall@k, Precision@k, MRR, nDCG |
| Generation | Is the answer faithful to the sources, correct, and on the point? | Faithfulness, answer relevance, correctness, citation accuracy |
The golden set
Each entry holds a question, a reference answer, and the passage or passages that contain it:
{
"question": "What is the deadline for submitting expense receipts?",
"reference_answer": "30 days from the date the expense was incurred.",
"relevant_chunks": ["finance-reimburse-v3#7"],
"type": "fact-lookup"
}
- Starting size: 50–200 questions is enough to separate one configuration from another.
- Take them from real questions — search logs, support tickets — not only from your own imagination.
- Mix the types: fact lookup, process explanation, comparison, and questions containing codes or proper nouns.
- Include questions with no answer in the corpus. Without them you cannot measure whether the system can say “I don't know”.
- An LLM may draft questions from the documents, but a person must approve them.
Retrieval metrics
Four questions, each with exactly one relevant passage — which is what makes the three formulas directly comparable here:
Question Rank of the correct passage Recall@5 1/rank 1/log2(rank+1)
Q1 1 1 1.0000 1.0000
Q2 3 1 0.3333 0.5000
Q3 not in the top 5 0 0.0000 0.0000
Q4 2 1 0.5000 0.6309
Recall@5 = 3/4 = 0.7500 MRR = 0.4583 nDCG@5 = 0.5327
The three disagree because they discount rank differently. At rank 2, 1/rank already halves the credit to 0.5000 while 1/log₂(rank+1) still awards 0.6309;
at rank 5 they are 0.2000 against 0.3869. MRR is therefore the harshest of the three about anything that is not first, which is why it moves the most when you add a reranker.
Answer metrics
- Faithfulness (groundedness): is every claim in the answer supported by the context? Low means the model is adding things.
- Answer relevance: does it address the question, or wander?
- Correctness: against the reference answer, is it right?
- Citation accuracy: does the cited passage actually contain the claim?
- Refusing at the right time: does it decline on unanswerable questions, and does it wrongly decline on answerable ones?
At scale these are usually scored by an LLM judge against an explicit rubric. Grade a small sample by hand first, to check the judge agrees with a person, before you trust the numbers. Libraries such as RAGAS, DeepEval and TruLens ship implementations of most of these.
The evaluation loop
- Freeze the golden set and the data version.
- Run the current configuration and keep the result as the baseline.
- Change one thing (chunk size, embedding model, reranker, prompt).
- Re-run, compare every metric, and look at which questions got worse — not only at the mean.
- Keep the change if it wins; put the evaluation in CI so nothing silently regresses.
A ranks the correct passage at 1, –, 2, 2. B ranks it at 2, 2, 1, –.
Both score Recall@5 = 0.75, MRR = 0.500, nDCG@5 = 0.5655 — all three aggregates identical to four decimal places. Yet A fails Q2 and B fails Q4. Whether that matters depends entirely on which question your users actually ask, and no average will ever tell you. Keep the per-question table, and diff it.
Monitoring in production
| Signal | What a bad trend means |
|---|---|
| Rate of “not found in the documents” | Rising: missing documents, a broken index, or new questions outside the scope |
| 👍/👎 and immediate re-asking | Users are not satisfied with the answer |
| Relevance score of the top-1 result | Drifting down: the corpus and the questions are diverging |
| p95 latency, tokens per question | Cost and experience |
| Index freshness | New documents have not been indexed yet |
Log the question, the retrieved passages and the answer, under whatever privacy controls apply. The 👎 cases are the best source of new golden-set entries you will ever get.