Practice

A miniature RAG laboratory

A complete RAG pipeline running in your browser over the staff handbook of a fictional company, “Sao Mai”: chunk → vectorise → retrieve top-k → build the prompt. Change a parameter and watch what moves.

About the word “embedding” here So that it runs with no server, this laboratory uses TF-IDF vectors — matching on shared words — in place of a real embedding model. It therefore handles questions phrased in the document's own words, and misses paraphrases. The fifth sample question shows that limit exactly; see Embeddings & vector stores and Retrieval for what fixes it.

Question

Parameters

The corpus

    Retrieval result

    The prompt sent to the LLM

    This is what the model actually reads. When nothing clears the threshold, the system should decline instead of calling the model.

    Experiments worth running

    Every number below was read off this page at the default settings (chunk 40, overlap 8, top-k 3, threshold 0.05, title prepended) unless stated otherwise.