Retrieval: hybrid search, reranking and query rewriting

Retrieval is where most of a RAG system's quality is decided. This lesson goes from a plain vector search to the multi-stage pipeline that production systems actually run.

Semantic search and keyword search

Dense (vectors)Sparse (BM25, keywords)
Strong atSynonyms, paraphrase, natural questionsCodes, proper nouns, rare terms, exact quotes
Weak at“SKU-48213”, “Clause 12.3”, in-house abbreviations“work from home” ≠ “remote working”
Typical toolspgvector, Qdrant, MilvusElasticsearch/OpenSearch, PostgreSQL full-text

BM25 scores a passage from how often the query's terms occur in it (with saturation), how rare those terms are across the corpus (IDF), and the passage length. It is an old, simple algorithm — and still a hard baseline to beat whenever the question contains the exact words.

The weakness in the second row, measured On PostgreSQL 16.15, eight short passages indexed with to_tsvector('english', …) and queried through websearch_to_tsquery — lexical matching with ts_rank_cd, the same family as BM25 though not identical to it. The query SKU-48213 returns exactly 1 row, the passage holding that code. The query remote working returns the remote-work policy. But the query work from home returns 0 rows — the answer is sitting in passage 1 and shares not one stem with the question. No parameter fixes this; only a second, semantic retriever does.

Hybrid search and Reciprocal Rank Fusion

Run both searches and merge the lists. Their scores are on different scales (cosine 0–1, BM25 unbounded), so the usual merge uses ranks rather than scores — Reciprocal Rank Fusion (RRF):

RRF(d) = Σlists 1 / (k + rank(d))   with k typically 60
                 vector   BM25    RRF (k = 60)
passage "Clause 12" rank 4   rank 1  1/64 + 1/61 = 0.0320
passage "WFH"       rank 1   rank 9  1/61 + 1/69 = 0.0309
passage "Leave"     rank 2   —       1/62        = 0.0161

A passage that does reasonably well in both lists rises to the top; one that is strong in only a single list still gets a chance. Note how small the margin is — 0.0320 against 0.0309, a gap of 3.7% — so with k = 60 a first place in one list does not simply overrule everything else.

Reranking

An embedding model encodes the question and the passage separately (a bi-encoder): fast, but coarse. A cross-encoder reads the question and the passage together, which judges relevance far more accurately — and is far too slow to run over a whole corpus.

Stage 1Hybrid search fetches 30–50 candidates
→
Stage 2Cross-encoder rescores each pair
→
ResultKeep the best 3–8 for the prompt

You can use an open-weights reranker or a provider's rerank API. Either way you are paying latency — commonly tens to hundreds of milliseconds — so measure whether the quality gain is worth it on your own questions.

Choosing top-k and a threshold

Transforming the question

Real questions are short, ambiguous, or depend on what was said earlier in the conversation. One small LLM step before retrieval can change the outcome substantially:

TechniqueExample
Conversational rewrite“And on probation?” → “Can an employee on probation work remotely?”
Multi-queryGenerate three phrasings, retrieve for each, merge with RRF
Decomposition“Compare the 2023 and 2025 leave policies” → two separate queries
HyDEHave the LLM write a hypothetical answer and embed that — it looks more like a document than the question does
Filter extraction“latest HR policy” → filter department = 'HR', order by updated_at
Every LLM step costs latency and money Start simple (hybrid plus rerank), measure, and only then add query transformation — for the specific kinds of question that are failing.

Metadata filters and access control

Permission checks belong inside the retrieval query, not after the answer has been written:

-- User belongs to groups {'sales', 'all-staff'}
SELECT content
FROM   chunks
WHERE  allowed_groups && ARRAY['sales', 'all-staff']   -- overlaps the user's groups
ORDER  BY embedding <=> $1
LIMIT  30;
Leaking through the answer Once a confidential passage is in the prompt, there is no reliable way to stop the model mentioning it. A passage the user may not read must never be retrieved in the first place.

Check yourself

Why does RRF use ranks instead of adding the cosine score to the BM25 score? The two scores live on different scales and have different distributions; adding them lets one side dominate arbitrarily. Ranks are comparable across systems.
Why not run the cross-encoder over the whole corpus from the start? A cross-encoder runs the model once per (question, passage) pair. Over millions of passages that is hopeless. It is for rescoring a few dozen candidates.
No passage clears the relevance threshold. What should the system do? Say the information was not found in the documents — optionally suggesting a rephrasing or a person to ask — rather than calling the LLM with empty or noisy context.