A short course on RAG

Retrieval-Augmented Generation,
one step at a time

A large language model writes fluently but has never seen your internal documents, and it will invent an answer rather than admit that. RAG addresses this by finding the relevant passages first and only then asking the model to answer from those passages. This course walks the whole pipeline, from splitting documents to measuring whether the answers are any good.

A RAG system, end to end

Indexing phase — runs ahead of time, periodically

Collect documents
→
Clean & attach metadata
→
Split into chunks
→
Compute embeddings
→
Store in a vector store

Query phase — runs for every question

Question
→
Fetch the top-k passages
→
Rerank
→
Build a prompt with sources
→
Model answers & cites

Course path

Each lesson is a 10–15 minute read and ends with questions to check yourself.

By the end you will be able to

  • Explain what problem RAG solves, and when it is the wrong tool.
  • Design the indexing phase: chunking, choice of embedding model, choice of vector store.
  • Combine keyword search with semantic search, then rerank for precision.
  • Write a prompt that forces the model to rely on its sources and to say "I don't know".
  • Measure retrieval and answer quality with numbers rather than impressions.
  • Diagnose the failures that show up once RAG meets a real product.

Reference material