Practice
A miniature RAG laboratory
A complete RAG pipeline running in your browser over the staff handbook of a fictional company, “Sao Mai”: chunk → vectorise → retrieve top-k → build the prompt. Change a parameter and watch what moves.
About the word “embedding” here
So that it runs with no server, this laboratory uses TF-IDF vectors — matching on shared words — in place of a real embedding model.
It therefore handles questions phrased in the document's own words, and misses paraphrases. The fifth sample question shows that limit exactly;
see Embeddings & vector stores and Retrieval for what fixes it.
Question
Parameters
The corpus
Retrieval result
The prompt sent to the LLM
This is what the model actually reads. When nothing clears the threshold, the system should decline instead of calling the model.
Experiments worth running
Every number below was read off this page at the default settings (chunk 40, overlap 8, top-k 3, threshold 0.05, title prepended) unless stated otherwise.
-
The default question. It returns
hr-remote#1at cosine 0.242 — but read that passage: it contains “completed probation” and stops before the sentence about the line manager, which sits inhr-remote#2at only 0.097. The top passage alone does not answer the question. This is what a chunk boundary costs you. - Drop the chunk size to 15. The corpus goes from 18 chunks to 64. The top score rises to 0.283 because each chunk is shorter and therefore more concentrated — and carries correspondingly less context around the matching words.
- Ask the paraphrase “Am I allowed to do my job from my own house?”. The remote-work passages score exactly 0.000: they share no term with the question at all. The best score anywhere in the corpus is 0.044, below the default threshold of 0.05, so the laboratory refuses to answer. Refusing is the honest outcome here — and a real embedding model would have found the passage.
-
Turn the title off. The direction of the change depends on the question, which is worth seeing for yourself.
For the six sample questions, prepending the title moves the top score by less than 0.04 and usually downwards, because it lengthens the vector with words the question never uses.
But ask
annual leave entitlementand the title is worth +0.134 (0.412 against 0.278) and even changes which chunk ranks first; askinformation securityand it is worth +0.124 (0.298 against 0.174). The title earns its place when the question echoes the document's subject, not otherwise. - Raise top-k from 3 to 8. Nothing happens: the context stays at 139 words, because the passages ranked 4 to 8 all score below the default threshold and are dropped before the prompt is built. Pull the threshold down to 0 as well and the context jumps to 355 words, about 2.6×. So it is the threshold, not k, that decides how much text the model ends up reading — and k only sets the ceiling. (The same experiment on the Vietnamese corpus behaves differently: there all eight passages clear 0.05, and the context does grow. Same parameters, different data, different outcome.)
- Ask something the handbook does not cover (“Does the company provide parking?”). The best score is 0.034, so the default threshold already rejects everything. Note that the system still ranks the passages — a ranking always exists. Only a threshold turns “the best of a bad set” into “I don't know”.