Retrieval: hybrid search, reranking and query rewriting
Retrieval is where most of a RAG system's quality is decided. This lesson goes from a plain vector search to the multi-stage pipeline that production systems actually run.
Semantic search and keyword search
| Dense (vectors) | Sparse (BM25, keywords) | |
|---|---|---|
| Strong at | Synonyms, paraphrase, natural questions | Codes, proper nouns, rare terms, exact quotes |
| Weak at | “SKU-48213”, “Clause 12.3”, in-house abbreviations | “work from home” ≠ “remote working” |
| Typical tools | pgvector, Qdrant, Milvus | Elasticsearch/OpenSearch, PostgreSQL full-text |
BM25 scores a passage from how often the query's terms occur in it (with saturation), how rare those terms are across the corpus (IDF), and the passage length. It is an old, simple algorithm — and still a hard baseline to beat whenever the question contains the exact words.
to_tsvector('english', …) and queried through websearch_to_tsquery — lexical matching with ts_rank_cd, the same family as BM25 though not identical to it.
The query SKU-48213 returns exactly 1 row, the passage holding that code. The query remote working returns the remote-work policy.
But the query work from home returns 0 rows — the answer is sitting in passage 1 and shares not one stem with the question.
No parameter fixes this; only a second, semantic retriever does.
Hybrid search and Reciprocal Rank Fusion
Run both searches and merge the lists. Their scores are on different scales (cosine 0–1, BM25 unbounded), so the usual merge uses ranks rather than scores — Reciprocal Rank Fusion (RRF):
vector BM25 RRF (k = 60)
passage "Clause 12" rank 4 rank 1 1/64 + 1/61 = 0.0320
passage "WFH" rank 1 rank 9 1/61 + 1/69 = 0.0309
passage "Leave" rank 2 — 1/62 = 0.0161
A passage that does reasonably well in both lists rises to the top; one that is strong in only a single list still gets a chance. Note how small the margin is — 0.0320 against 0.0309, a gap of 3.7% — so with k = 60 a first place in one list does not simply overrule everything else.
Reranking
An embedding model encodes the question and the passage separately (a bi-encoder): fast, but coarse. A cross-encoder reads the question and the passage together, which judges relevance far more accurately — and is far too slow to run over a whole corpus.
You can use an open-weights reranker or a provider's rerank API. Either way you are paying latency — commonly tens to hundreds of milliseconds — so measure whether the quality gain is worth it on your own questions.
Choosing top-k and a threshold
- k too small: the passage holding the answer gets missed — low recall.
- k too large: a long, expensive prompt, and irrelevant passages that can pull the model off course.
- A score threshold: drop passages below a relevance floor. If nothing survives, answer “not found in the documents” rather than letting the model improvise.
- Diversity (MMR): stop five near-identical passages from occupying every slot.
Transforming the question
Real questions are short, ambiguous, or depend on what was said earlier in the conversation. One small LLM step before retrieval can change the outcome substantially:
| Technique | Example |
|---|---|
| Conversational rewrite | “And on probation?” → “Can an employee on probation work remotely?” |
| Multi-query | Generate three phrasings, retrieve for each, merge with RRF |
| Decomposition | “Compare the 2023 and 2025 leave policies” → two separate queries |
| HyDE | Have the LLM write a hypothetical answer and embed that — it looks more like a document than the question does |
| Filter extraction | “latest HR policy” → filter department = 'HR', order by updated_at |
Metadata filters and access control
Permission checks belong inside the retrieval query, not after the answer has been written:
-- User belongs to groups {'sales', 'all-staff'}
SELECT content
FROM chunks
WHERE allowed_groups && ARRAY['sales', 'all-staff'] -- overlaps the user's groups
ORDER BY embedding <=> $1
LIMIT 30;