A short course on RAG
Retrieval-Augmented Generation,
one step at a time
A large language model writes fluently but has never seen your internal documents, and it will invent an answer rather than admit that. RAG addresses this by finding the relevant passages first and only then asking the model to answer from those passages. This course walks the whole pipeline, from splitting documents to measuring whether the answers are any good.
A RAG system, end to end
Indexing phase — runs ahead of time, periodically
Query phase — runs for every question
Course path
Each lesson is a 10–15 minute read and ends with questions to check yourself.
What RAG is, and when you need it
The limits of an LLM, the idea behind RAG, and how it compares with fine-tuning and long context.
DataPreparing data & chunking
Cleaning, metadata, splitting strategies and how to choose a chunk size.
VectorsEmbeddings & vector stores
Semantic vectors, cosine similarity, ANN indexes, and a pgvector example.
RetrievalRetrieval: hybrid, reranking
BM25 plus vectors, Reciprocal Rank Fusion, cross-encoders, query rewriting.
AnsweringBuilding the prompt & answering
Prompt templates, citing sources, the token budget, resisting prompt injection.
EvaluationMeasuring quality
Recall@k, MRR, faithfulness, a golden question set, and monitoring in production.
OperationsCommon failures & what is next
A symptom → cause → fix table, plus GraphRAG and agentic RAG.
PracticeRAG laboratory
Change the chunk size and top-k, then watch what retrieval returns and how the prompt shifts.
By the end you will be able to
- Explain what problem RAG solves, and when it is the wrong tool.
- Design the indexing phase: chunking, choice of embedding model, choice of vector store.
- Combine keyword search with semantic search, then rerank for precision.
- Write a prompt that forces the model to rely on its sources and to say "I don't know".
- Measure retrieval and answer quality with numbers rather than impressions.
- Diagnose the failures that show up once RAG meets a real product.