Measuring RAG quality

“I tried a few questions and it seemed fine” is not an evaluation. This lesson builds a measurable, repeatable process, so that every change to chunking, models or prompts arrives with a number attached.

Two layers to measure

Measure them separately, or you will not know which one broke:

LayerThe question it answersMetrics
RetrievalWas the passage holding the answer fetched, and fetched high up?Recall@k, Precision@k, MRR, nDCG
GenerationIs the answer faithful to the sources, correct, and on the point?Faithfulness, answer relevance, correctness, citation accuracy

The golden set

Each entry holds a question, a reference answer, and the passage or passages that contain it:

{
  "question": "What is the deadline for submitting expense receipts?",
  "reference_answer": "30 days from the date the expense was incurred.",
  "relevant_chunks": ["finance-reimburse-v3#7"],
  "type": "fact-lookup"
}

Retrieval metrics

Recall@k = relevant passages within the top k / total relevant passages
MRR = mean over questions of 1 / (rank of the first relevant passage)
nDCG@k = mean over questions of 1 / log2(rank + 1)  (binary relevance, one relevant passage)

Four questions, each with exactly one relevant passage — which is what makes the three formulas directly comparable here:

Question   Rank of the correct passage   Recall@5   1/rank   1/log2(rank+1)
Q1         1                             1          1.0000   1.0000
Q2         3                             1          0.3333   0.5000
Q3         not in the top 5              0          0.0000   0.0000
Q4         2                             1          0.5000   0.6309

Recall@5 = 3/4 = 0.7500     MRR = 0.4583     nDCG@5 = 0.5327

The three disagree because they discount rank differently. At rank 2, 1/rank already halves the credit to 0.5000 while 1/log₂(rank+1) still awards 0.6309; at rank 5 they are 0.2000 against 0.3869. MRR is therefore the harshest of the three about anything that is not first, which is why it moves the most when you add a reranker.

Recall first, precision second For RAG, recall@k in the retrieval layer matters most: if the right passage never arrives, no prompt can rescue the answer. Reranking is what raises precision afterwards.

Answer metrics

At scale these are usually scored by an LLM judge against an explicit rubric. Grade a small sample by hand first, to check the judge agrees with a person, before you trust the numbers. Libraries such as RAGAS, DeepEval and TruLens ship implementations of most of these.

The evaluation loop

  1. Freeze the golden set and the data version.
  2. Run the current configuration and keep the result as the baseline.
  3. Change one thing (chunk size, embedding model, reranker, prompt).
  4. Re-run, compare every metric, and look at which questions got worse — not only at the mean.
  5. Keep the change if it wins; put the evaluation in CI so nothing silently regresses.
Why step 4 says “not only at the mean” Two configurations, same four questions, one relevant passage each:
A ranks the correct passage at 1, –, 2, 2.   B ranks it at 2, 2, 1, –.
Both score Recall@5 = 0.75, MRR = 0.500, nDCG@5 = 0.5655 — all three aggregates identical to four decimal places. Yet A fails Q2 and B fails Q4. Whether that matters depends entirely on which question your users actually ask, and no average will ever tell you. Keep the per-question table, and diff it.

Monitoring in production

SignalWhat a bad trend means
Rate of “not found in the documents”Rising: missing documents, a broken index, or new questions outside the scope
👍/👎 and immediate re-askingUsers are not satisfied with the answer
Relevance score of the top-1 resultDrifting down: the corpus and the questions are diverging
p95 latency, tokens per questionCost and experience
Index freshnessNew documents have not been indexed yet

Log the question, the retrieved passages and the answer, under whatever privacy controls apply. The 👎 cases are the best source of new golden-set entries you will ever get.

Check yourself

High faithfulness but low correctness — what does that mean? The model is faithful to its context, but the context is wrong or incomplete. The fault is in retrieval or in the data itself (an outdated document, say), not in generation.
Recall@5 is 0.95 and users still report wrong answers. What do you check next? The generation layer: faithfulness, the prompt, the order of the context, whether the model is using the right passage. Also check that the golden set actually resembles the questions being asked.
Why does the golden set need questions with no answer? To measure — and prevent — invention when information is missing. Without them, a system that confidently answers everything can still post an excellent score.