Embeddings & vector stores
An embedding turns text into coordinates in a high-dimensional space, where “close in meaning” becomes “close in distance”. A vector store is what finds those close points quickly.
What an embedding is
An embedding model takes a passage and returns a vector — typically a few hundred to a few thousand real numbers. It is trained so that passages with similar meaning land near each other, even when they share no words at all:
| Question | Lands near |
|---|---|
| “I'd like to work from home a few days” | “Remote work policy” |
| “Who do I tell if my laptop breaks?” | “Equipment incident support process” |
| “How many days off do I get a year?” | “Annual leave entitlement” |
This is the advantage over pure keyword search — and also the reason this course's laboratory, which uses TF-IDF and therefore relies on shared words, will miss every pair in the table above.
Measuring closeness: cosine similarity
The usual measure is the cosine of the angle between two vectors:
A value near 1 means the same direction (very similar); near 0 means unrelated. With three dimensions, so it can be checked by hand:
question q = [0.9, 0.1, 0.3]
passage A a = [0.8, 0.2, 0.4] (about annual leave)
passage B b = [0.1, 0.9, 0.2] (about equipment)
q · a = 0.72 + 0.02 + 0.12 = 0.86 ‖q‖ ≈ 0.9539 ‖a‖ ≈ 0.9165
cos(q, a) = 0.86 / (0.9539 × 0.9165) = 0.9836
q · b = 0.09 + 0.09 + 0.06 = 0.24 ‖b‖ ≈ 0.9274
cos(q, b) = 0.24 / (0.9539 × 0.9274) = 0.2713
Passage A ranks above passage B, by a factor of 3.6. Many systems normalise vectors to length 1; the cosine is then exactly the dot product, which is faster to compute.
Choosing an embedding model
- Language. If your corpus is not in English, pick a multilingual model and test it on your own questions — the gap between models is large and does not track their English benchmark scores.
- One model for both sides. Vectors from two different models are not comparable. Changing model means re-embedding the entire corpus.
- Dimensionality. More dimensions are usually more accurate but cost memory and time. Some models support truncating the vector to fewer dimensions.
- Input token limit. A chunk longer than the limit is truncated, sometimes without warning.
- Query and document prefixes. Several models require a different prefix for questions and for passages. Read the model card before you index a million chunks.
Vector stores and ANN indexes
With a few thousand passages you can compare the question against every vector (brute force). With millions you need an ANN — Approximate Nearest Neighbor index: it accepts occasionally missing a neighbour in exchange for being orders of magnitude faster.
| Index | Idea | Trade-off |
|---|---|---|
| HNSW | A layered graph; descend from sparse layers to dense ones towards the query | Fast, high recall; heavy on RAM, slower to build |
| IVF | Cluster the vectors, then search only the nearest few clusters | Lighter; needs a training step, recall depends on how many clusters are probed |
| PQ (quantisation) | Compress each vector into a short code | Large memory savings; loses precision |
Common choices: pgvector (if you already run PostgreSQL), Qdrant, Milvus, Weaviate, Elasticsearch/OpenSearch (strong at keyword search too), or the FAISS library if you are managing it yourself.
A PostgreSQL + pgvector example
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (
id bigserial PRIMARY KEY,
doc_id text NOT NULL,
section text,
content text NOT NULL,
department text,
updated_at date,
embedding vector(1024) -- exactly the model's dimensionality
);
-- HNSW index on cosine distance
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
-- Top-5 passages nearest the question, restricted to HR documents
-- <=> is cosine DISTANCE, so similarity = 1 - distance
SELECT doc_id, section, content,
1 - (embedding <=> $1) AS similarity
FROM chunks
WHERE department = 'HR'
ORDER BY embedding <=> $1
LIMIT 5;
department = 'HR', an HNSW index on vector_cosine_ops, and the query above with LIMIT 5.
At the default hnsw.ef_search = 40 it returns 3 rows, not 5. EXPLAIN ANALYZE shows why:
Index Scan using chunks_hnsw (actual rows=3) with Rows Removed by Filter: 388 — the index produced 391 candidates by distance alone, and the WHERE clause then discarded all but three.
Raising the setting to SET hnsw.ef_search = 100 returns all 5, as does an exact scan. Nothing warns you; the query simply succeeds with a short result.
So: count the rows you actually received, raise the search parameter, or use a store that filters before the index search rather than after.
Check yourself
You switch to a better embedding model. What has to happen to the existing data?
Re-embed the whole corpus with the new model, and change the column's dimensionality to match. Vectors from two models must never share one index.Why “approximate” nearest neighbor?
The index trades exactness for speed: it does not guarantee the true nearest neighbours. How far it goes is controlled by a parameter — in pgvector's HNSW,hnsw.ef_search, the number of candidates it keeps while descending the graph.