Common failures & what comes next

A diagnostic table for when a RAG system misbehaves, and the directions worth taking once plain RAG stops being enough.

Diagnosing in order

For every wrong answer, walk the pipeline in sequence and stop at the first “no”:

  1. Do the documents contain the answer at all? If not, this is a data problem, not a RAG problem.
  2. Does the answer sit whole inside one chunk? If it is cut in half or stripped of its context — fix chunking.
  3. Was that chunk retrieved, inside the top k? If not — fix retrieval: hybrid search, the embedding model, query transformation.
  4. Did the chunk reach the prompt? A reranker may have dropped it, a threshold may have cut it, or the token budget may have truncated it.
  5. It was in the prompt and the answer is still wrong? Fix the prompt, the context order, or change the generating model.
You cannot diagnose what you did not log Store the original question, the transformed question, the passage list with its score at every stage, the final prompt, and the answer. Without that trail you are guessing.

Symptom → cause → fix

SymptomUsual causeFix
Says “I don't know” although the document exists Retrieval missed it; the question words differ from the document's; a code or proper noun Hybrid search, prepend the title to the chunk, multi-query, check the threshold is not set too high
Wrong, but extremely convincing A near-miss passage; the model inferring past its sources Rerank, lower k, force citations in the prompt, measure faithfulness
Cites a superseded document Several versions coexist; the index is stale Version and effective-date metadata, filter to the latest, re-index when a document changes
Half an answer The information spans several chunks; a table was cut Structural chunking, parent–child, merge adjacent passages, raise k deliberately
A user sees something they are not allowed to No permission filter at retrieval time; a shared cache ACL filter inside the vector-store query, permission-scoped cache keys, permission tests in CI
“Compare”, “summarise”, “list every…” answered badly Top-k passages cannot cover the whole corpus Decompose the question, multi-step retrieval, structured data plus SQL, or GraphRAG
Slow and expensive Large k, long context, several auxiliary LLM steps Rerank then cut k, a token budget, caching, a small model for side steps, streaming
Noticeably worse in one language than another The embedding or rerank model is weak in that language; Unicode normalisation differs between the query and the index Pick a multilingual model and measure it on that language's own questions; normalise to NFC on both sides; use a tokeniser that suits the language for BM25
Strange behaviour around particular documents Prompt injection sitting in retrieved content Wrap the context, instruct that data is not commands, control what is indexed, limit what the model may act on
Sources contradict each other Departments have not reconciled their documents Prefer the authoritative or newer source through metadata; let the model state the contradiction and cite both

Beyond basic RAG

Contextual retrieval

Before embedding, have an LLM write one sentence of context for each chunk (“This passage is section X of document Y, and it describes…”) and prepend it. You pay once, at indexing time, and get better retrieval for passages that are opaque on their own.

Agentic RAG

Instead of a single “retrieve then answer” pass, the LLM decides for itself: whether to search, what to search for, which source to use (documents, SQL, an API), reads the result and searches again. Strong on multi-step questions, but harder to keep within a latency and cost budget, and harder to evaluate.

GraphRAG

Extract entities and relations from the documents into a knowledge graph, with summaries per topic cluster. Useful for corpus-wide questions (“what are the main themes in customer feedback?”) that top-k passages simply cannot answer.

Structured data

For numbers that live in a database, having the LLM generate SQL — read-only, against a restricted set of tables — is usually far more accurate than embedding tables of figures.

Do not skip ahead Most systems get most of their quality from clean data, sensible chunking, hybrid search, a reranker and a good prompt. Add the advanced techniques only when your evaluation set shows basic RAG has hit its ceiling.

The course in four lines

Data

Clean, structured, with metadata

Input quality sets the ceiling for everything downstream.

Retrieval

Hybrid + rerank + permission filter

Where most right and wrong answers are decided.

Generation

Grounded, cited, willing to refuse

A prompt with a clear line between data and instructions.

Evaluation

Measure each layer, every change

A golden set, plus production monitoring.

Check yourself

A user asks “List every policy effective from 2025”. Why does top-k RAG answer incompletely? The question requires scanning the whole corpus, while top-k returns only the few nearest passages. Use a metadata filter on the effective date to build the list, or query structured data.
The right passage is in the top 30 but never reaches the prompt. Which step do you suspect? The reranker dropped it, the score threshold was too high, or the token budget truncated it. Look up that passage's score at each stage in the log.