Embeddings & vector stores

An embedding turns text into coordinates in a high-dimensional space, where “close in meaning” becomes “close in distance”. A vector store is what finds those close points quickly.

What an embedding is

An embedding model takes a passage and returns a vector — typically a few hundred to a few thousand real numbers. It is trained so that passages with similar meaning land near each other, even when they share no words at all:

QuestionLands near
“I'd like to work from home a few days”“Remote work policy”
“Who do I tell if my laptop breaks?”“Equipment incident support process”
“How many days off do I get a year?”“Annual leave entitlement”

This is the advantage over pure keyword search — and also the reason this course's laboratory, which uses TF-IDF and therefore relies on shared words, will miss every pair in the table above.

Measuring closeness: cosine similarity

The usual measure is the cosine of the angle between two vectors:

cos(a, b) = (a · b) / (‖a‖ × ‖b‖)

A value near 1 means the same direction (very similar); near 0 means unrelated. With three dimensions, so it can be checked by hand:

question   q = [0.9, 0.1, 0.3]
passage A  a = [0.8, 0.2, 0.4]   (about annual leave)
passage B  b = [0.1, 0.9, 0.2]   (about equipment)

q · a = 0.72 + 0.02 + 0.12 = 0.86      ‖q‖ ≈ 0.9539   ‖a‖ ≈ 0.9165
cos(q, a) = 0.86 / (0.9539 × 0.9165) = 0.9836

q · b = 0.09 + 0.09 + 0.06 = 0.24      ‖b‖ ≈ 0.9274
cos(q, b) = 0.24 / (0.9539 × 0.9274) = 0.2713

Passage A ranks above passage B, by a factor of 3.6. Many systems normalise vectors to length 1; the cosine is then exactly the dot product, which is faster to compute.

Choosing an embedding model

Vector stores and ANN indexes

With a few thousand passages you can compare the question against every vector (brute force). With millions you need an ANN — Approximate Nearest Neighbor index: it accepts occasionally missing a neighbour in exchange for being orders of magnitude faster.

IndexIdeaTrade-off
HNSWA layered graph; descend from sparse layers to dense ones towards the queryFast, high recall; heavy on RAM, slower to build
IVFCluster the vectors, then search only the nearest few clustersLighter; needs a training step, recall depends on how many clusters are probed
PQ (quantisation)Compress each vector into a short codeLarge memory savings; loses precision

Common choices: pgvector (if you already run PostgreSQL), Qdrant, Milvus, Weaviate, Elasticsearch/OpenSearch (strong at keyword search too), or the FAISS library if you are managing it yourself.

The criteria that actually decide it Metadata filtering inside the same query, hybrid search, updating and deleting individual chunks, backups, and whether your team can operate the thing. Below a few million chunks, PostgreSQL + pgvector is usually the least trouble.

A PostgreSQL + pgvector example

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE chunks (
  id          bigserial PRIMARY KEY,
  doc_id      text NOT NULL,
  section     text,
  content     text NOT NULL,
  department  text,
  updated_at  date,
  embedding   vector(1024)          -- exactly the model's dimensionality
);

-- HNSW index on cosine distance
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

-- Top-5 passages nearest the question, restricted to HR documents
-- <=> is cosine DISTANCE, so similarity = 1 - distance
SELECT doc_id, section, content,
       1 - (embedding <=> $1) AS similarity
FROM   chunks
WHERE  department = 'HR'
ORDER  BY embedding <=> $1
LIMIT  5;
A selective filter can silently return fewer rows than you asked for Measured on PostgreSQL 16.15 with pgvector 0.8.6: 50,000 chunks of 128 dimensions, of which 500 (1%) carry department = 'HR', an HNSW index on vector_cosine_ops, and the query above with LIMIT 5. At the default hnsw.ef_search = 40 it returns 3 rows, not 5. EXPLAIN ANALYZE shows why: Index Scan using chunks_hnsw (actual rows=3) with Rows Removed by Filter: 388 — the index produced 391 candidates by distance alone, and the WHERE clause then discarded all but three. Raising the setting to SET hnsw.ef_search = 100 returns all 5, as does an exact scan. Nothing warns you; the query simply succeeds with a short result. So: count the rows you actually received, raise the search parameter, or use a store that filters before the index search rather than after.

Check yourself

You switch to a better embedding model. What has to happen to the existing data? Re-embed the whole corpus with the new model, and change the column's dimensionality to match. Vectors from two models must never share one index.
Why “approximate” nearest neighbor? The index trades exactness for speed: it does not guarantee the true nearest neighbours. How far it goes is controlled by a parameter — in pgvector's HNSW, hnsw.ef_search, the number of candidates it keeps while descending the graph.
A question contains the product code “SKU-48213”. Will embeddings find it? Usually not: codes, part numbers and rare proper nouns are where embeddings are weakest. This is precisely why keyword search (BM25) is combined with them — see Retrieval.