Reference
RAG glossary
Short definitions of the terms used across this course, each pointing at the lesson that develops it.
| Term | Meaning | Lesson |
|---|---|---|
| RAG (Retrieval-Augmented Generation) | Retrieve the relevant passages, put them in the prompt, and have the LLM answer from them. | 1 |
| Hallucination | The model producing something plausible-sounding that is wrong or unsupported. | 1 |
| Indexing | The preparation phase: load, clean, chunk, embed and store the documents. | 1 |
| Chunk | A small piece of a document — the unit that gets embedded, retrieved and put in the prompt. | 2 |
| Overlap | Content repeated between two adjacent chunks, so nothing on a boundary is lost. | 2 |
| Metadata | What travels with a chunk: source, section, version, date, department, permission groups. | 2 |
| Parent–child (small-to-big) | Retrieve on the small chunk, put the larger parent passage in the prompt. | 2 |
| Token | The unit of text a model processes; both the context limit and the bill are counted in tokens. | 2 |
| Embedding | A numeric vector standing for the meaning of a text; similar meanings land near each other. | 3 |
| Cosine similarity | Similarity as the angle between two vectors: a·b / (‖a‖‖b‖). | 3 |
| Vector store | A system that stores vectors and searches them by proximity, usually with metadata filtering. | 3 |
| ANN (Approximate Nearest Neighbor) | Finding near neighbours approximately, trading a little accuracy for a lot of speed. | 3 |
| HNSW | A layered-graph ANN index; fast, high recall, memory hungry. | 3 |
| ef_search | How many candidates HNSW keeps while descending the graph. Too low, and a selective filter returns fewer rows than you asked for. | 3 |
| Dense retrieval | Retrieval through embeddings — by meaning. | 4 |
| Sparse retrieval / BM25 | Keyword retrieval, scored by term frequency and term rarity. | 4 |
| TF-IDF | Term weight = frequency in the passage × rarity across the corpus; the basis of classical keyword search. | Lab |
| Hybrid search | Running dense and sparse retrieval and merging the results. | 4 |
| RRF (Reciprocal Rank Fusion) | Merging several ranked lists with Σ 1/(k + rank). | 4 |
| Top-k | How many highest-scoring passages survive retrieval. | 4 |
| Reranker / cross-encoder | A model that reads question and passage together to rescore relevance — more accurate, much slower. | 4 |
| MMR (Maximal Marginal Relevance) | Picking passages that are relevant and different from those already picked. | 4 |
| HyDE | Generating a hypothetical answer and retrieving with that instead of the question. | 4 |
| Multi-query | Generating several phrasings of the question, retrieving for each, then merging. | 4 |
| Context window | The maximum number of tokens a model can read in one call. | 5 |
| Lost in the middle | The tendency to under-use information buried in the middle of a long context. | 5 |
| Grounding / citation | Tying every claim in the answer to a specific source in the context. | 5 |
| Prompt injection | Text — possibly inside a retrieved document — that tries to make the model disobey its instructions. | 5 |
| Golden set | Questions with known answers and known relevant passages, used for evaluation. | 6 |
| Recall@k | The share of relevant passages that appear within the top k results. | 6 |
| MRR | The mean reciprocal rank of the first relevant passage. | 6 |
| nDCG | A ranking metric that accounts for both degree of relevance and position, normalised to 0–1. | 6 |
| Faithfulness | How far the answer is supported by the context, with nothing added. | 6 |
| LLM-as-judge | Scoring answers with an LLM against a rubric; must be checked against human scoring. | 6 |
| Contextual retrieval | Prepending an LLM-written context sentence to each chunk before embedding it. | 7 |
| Agentic RAG | The LLM planning multi-step retrieval for itself, choosing tools and sources. | 7 |
| GraphRAG | Building an entity–relation graph from the documents to answer corpus-wide questions. | 7 |