Reference

RAG glossary

Short definitions of the terms used across this course, each pointing at the lesson that develops it.

TermMeaningLesson
RAG (Retrieval-Augmented Generation)Retrieve the relevant passages, put them in the prompt, and have the LLM answer from them.1
HallucinationThe model producing something plausible-sounding that is wrong or unsupported.1
IndexingThe preparation phase: load, clean, chunk, embed and store the documents.1
ChunkA small piece of a document — the unit that gets embedded, retrieved and put in the prompt.2
OverlapContent repeated between two adjacent chunks, so nothing on a boundary is lost.2
MetadataWhat travels with a chunk: source, section, version, date, department, permission groups.2
Parent–child (small-to-big)Retrieve on the small chunk, put the larger parent passage in the prompt.2
TokenThe unit of text a model processes; both the context limit and the bill are counted in tokens.2
EmbeddingA numeric vector standing for the meaning of a text; similar meanings land near each other.3
Cosine similaritySimilarity as the angle between two vectors: a·b / (‖a‖‖b‖).3
Vector storeA system that stores vectors and searches them by proximity, usually with metadata filtering.3
ANN (Approximate Nearest Neighbor)Finding near neighbours approximately, trading a little accuracy for a lot of speed.3
HNSWA layered-graph ANN index; fast, high recall, memory hungry.3
ef_searchHow many candidates HNSW keeps while descending the graph. Too low, and a selective filter returns fewer rows than you asked for.3
Dense retrievalRetrieval through embeddings — by meaning.4
Sparse retrieval / BM25Keyword retrieval, scored by term frequency and term rarity.4
TF-IDFTerm weight = frequency in the passage × rarity across the corpus; the basis of classical keyword search.Lab
Hybrid searchRunning dense and sparse retrieval and merging the results.4
RRF (Reciprocal Rank Fusion)Merging several ranked lists with Σ 1/(k + rank).4
Top-kHow many highest-scoring passages survive retrieval.4
Reranker / cross-encoderA model that reads question and passage together to rescore relevance — more accurate, much slower.4
MMR (Maximal Marginal Relevance)Picking passages that are relevant and different from those already picked.4
HyDEGenerating a hypothetical answer and retrieving with that instead of the question.4
Multi-queryGenerating several phrasings of the question, retrieving for each, then merging.4
Context windowThe maximum number of tokens a model can read in one call.5
Lost in the middleThe tendency to under-use information buried in the middle of a long context.5
Grounding / citationTying every claim in the answer to a specific source in the context.5
Prompt injectionText — possibly inside a retrieved document — that tries to make the model disobey its instructions.5
Golden setQuestions with known answers and known relevant passages, used for evaluation.6
Recall@kThe share of relevant passages that appear within the top k results.6
MRRThe mean reciprocal rank of the first relevant passage.6
nDCGA ranking metric that accounts for both degree of relevance and position, normalised to 0–1.6
FaithfulnessHow far the answer is supported by the context, with nothing added.6
LLM-as-judgeScoring answers with an LLM against a rubric; must be checked against human scoring.6
Contextual retrievalPrepending an LLM-written context sentence to each chunk before embedding it.7
Agentic RAGThe LLM planning multi-step retrieval for itself, choosing tools and sources.7
GraphRAGBuilding an entity–relation graph from the documents to answer corpus-wide questions.7