Preparing data & chunking

“Garbage in, garbage out” applies to RAG more literally than to most systems. How you clean and split your documents decides whether the right passage can be found at all.

Loading and cleaning

Real documents are rarely clean. Before splitting anything, deal with:

Tables are the classic weak point A price table extracted as a loose stream of numbers means nothing to an embedding model or to an LLM. Convert tables to Markdown, or to one sentence per row (“Pro plan: $29/month, 20 seats”), before you chunk.

Metadata, the part people forget

Every chunk should carry where it came from:

{
  "chunk_id": "hr-remote-v2.1#4",
  "doc_title": "Remote work policy",
  "section": "3. Who this applies to",
  "source_url": "https://intranet.example/hr/remote",
  "version": "2.1",
  "updated_at": "2025-03-01",
  "department": "HR",
  "allowed_groups": ["all-staff"]
}

Metadata does three jobs: filtering (search HR documents only, latest version only), access control (a user only ever retrieves chunks their groups allow) and citation (show the document title, the section, the link).

Why split at all

But passages that are too short break in the other direction: “The deadline is 30 days.” on its own does not say the deadline for what.

Chunking strategies

StrategyHow it worksSuitsWatch out for
Fixed sizeCut every N tokens, overlapping MA starting point; uniform textCuts mid-sentence and mid-table
RecursiveTry paragraph → sentence → word until it fitsMost proseA good default; what most libraries ship
StructuralSplit on headings, sections, clausesPolicies, technical docs, Markdown/HTMLAn over-long section still needs splitting
SemanticSplit where the embeddings of successive sentences divergeLong text that keeps changing subjectMore compute; measure before you believe it helps
Parent–child (small-to-big)Retrieve on small chunks, put the larger parent in the promptNeeding precise retrieval and full context at onceYou must store the parent–child relation

The trick worth knowing: put the context inside the chunk

Before embedding, prepend the document title and the section path to each chunk. “The deadline is 30 days.” becomes:

Expense reimbursement policy › 4. Deadline for submitting receipts
The deadline is 30 days from the date the expense was incurred.

One extra line, and both the retriever and the model now know what the passage is about.

Overlapping chunks, in code

def chunk_words(text, size=120, overlap=20):
    words = text.split()
    step = size - overlap
    chunks = []
    for start in range(0, len(words), step):
        piece = words[start:start + size]
        if piece:
            chunks.append(" ".join(piece))
        if start + size >= len(words):
            break
    return chunks

The overlap repeats the last few sentences of one chunk at the start of the next, so an idea sitting exactly on a boundary still appears whole in at least one chunk. Run on a 1,000-word document this returns 10 chunks: nine of 120 words and a last one of 100, with exactly 20 words shared between each neighbouring pair, and every word present in at least one chunk. (In production, count the tokens of your embedding model rather than whitespace-separated words.)

Choosing size and overlap

Do not guess — measure Write 30–50 questions whose answers you already know, try two or three chunk configurations, and compare recall@k (see Measuring quality). The best numbers depend on your own corpus. Try it now in the laboratory: drag the chunk size and watch the correct passage rise or fall in the ranking.

Check yourself

Why prepend the document title to a chunk's text before embedding it? Because many passages only mean something inside the section that contains them. The title steers the vector towards the right topic, which makes the chunk findable, and it tells the model what it is reading.
What is wrong with a 2,000-token chunk compared with a 300-token one? The vector is diluted across too many ideas, so it matches specific questions poorly; it burns context and money; and citations become vague. In exchange it loses less context — which is exactly why you measure, or use parent–child.
What is allowed_groups for, and at which step does it have to be applied? Access control: filter inside the query to the vector store, before any passage reaches the prompt. Filtering after the answer has been generated is far too late.