Common failures & what comes next
A diagnostic table for when a RAG system misbehaves, and the directions worth taking once plain RAG stops being enough.
Diagnosing in order
For every wrong answer, walk the pipeline in sequence and stop at the first “no”:
- Do the documents contain the answer at all? If not, this is a data problem, not a RAG problem.
- Does the answer sit whole inside one chunk? If it is cut in half or stripped of its context — fix chunking.
- Was that chunk retrieved, inside the top k? If not — fix retrieval: hybrid search, the embedding model, query transformation.
- Did the chunk reach the prompt? A reranker may have dropped it, a threshold may have cut it, or the token budget may have truncated it.
- It was in the prompt and the answer is still wrong? Fix the prompt, the context order, or change the generating model.
Symptom → cause → fix
| Symptom | Usual cause | Fix |
|---|---|---|
| Says “I don't know” although the document exists | Retrieval missed it; the question words differ from the document's; a code or proper noun | Hybrid search, prepend the title to the chunk, multi-query, check the threshold is not set too high |
| Wrong, but extremely convincing | A near-miss passage; the model inferring past its sources | Rerank, lower k, force citations in the prompt, measure faithfulness |
| Cites a superseded document | Several versions coexist; the index is stale | Version and effective-date metadata, filter to the latest, re-index when a document changes |
| Half an answer | The information spans several chunks; a table was cut | Structural chunking, parent–child, merge adjacent passages, raise k deliberately |
| A user sees something they are not allowed to | No permission filter at retrieval time; a shared cache | ACL filter inside the vector-store query, permission-scoped cache keys, permission tests in CI |
| “Compare”, “summarise”, “list every…” answered badly | Top-k passages cannot cover the whole corpus | Decompose the question, multi-step retrieval, structured data plus SQL, or GraphRAG |
| Slow and expensive | Large k, long context, several auxiliary LLM steps | Rerank then cut k, a token budget, caching, a small model for side steps, streaming |
| Noticeably worse in one language than another | The embedding or rerank model is weak in that language; Unicode normalisation differs between the query and the index | Pick a multilingual model and measure it on that language's own questions; normalise to NFC on both sides; use a tokeniser that suits the language for BM25 |
| Strange behaviour around particular documents | Prompt injection sitting in retrieved content | Wrap the context, instruct that data is not commands, control what is indexed, limit what the model may act on |
| Sources contradict each other | Departments have not reconciled their documents | Prefer the authoritative or newer source through metadata; let the model state the contradiction and cite both |
Beyond basic RAG
Contextual retrieval
Before embedding, have an LLM write one sentence of context for each chunk (“This passage is section X of document Y, and it describes…”) and prepend it. You pay once, at indexing time, and get better retrieval for passages that are opaque on their own.
Agentic RAG
Instead of a single “retrieve then answer” pass, the LLM decides for itself: whether to search, what to search for, which source to use (documents, SQL, an API), reads the result and searches again. Strong on multi-step questions, but harder to keep within a latency and cost budget, and harder to evaluate.
GraphRAG
Extract entities and relations from the documents into a knowledge graph, with summaries per topic cluster. Useful for corpus-wide questions (“what are the main themes in customer feedback?”) that top-k passages simply cannot answer.
Structured data
For numbers that live in a database, having the LLM generate SQL — read-only, against a restricted set of tables — is usually far more accurate than embedding tables of figures.
The course in four lines
Clean, structured, with metadata
Input quality sets the ceiling for everything downstream.
Hybrid + rerank + permission filter
Where most right and wrong answers are decided.
Grounded, cited, willing to refuse
A prompt with a clear line between data and instructions.
Measure each layer, every change
A golden set, plus production monitoring.