Preparing data & chunking
“Garbage in, garbage out” applies to RAG more literally than to most systems. How you clean and split your documents decides whether the right passage can be found at all.
Loading and cleaning
Real documents are rarely clean. Before splitting anything, deal with:
- Extracting text with its structure intact: keep headings, lists and tables. Scanned PDFs need OCR; multi-column PDFs interleave lines if the extractor is naive.
- Removing noise: running headers and footers, page numbers, site navigation, cookie banners, email signatures.
- Normalising: Unicode (precomposed and decomposed forms look identical but differ byte for byte — settle on NFC), whitespace, stray control characters.
- De-duplicating: several copies of one document make retrieval return the same passage repeatedly and crowd out everything else.
Metadata, the part people forget
Every chunk should carry where it came from:
{
"chunk_id": "hr-remote-v2.1#4",
"doc_title": "Remote work policy",
"section": "3. Who this applies to",
"source_url": "https://intranet.example/hr/remote",
"version": "2.1",
"updated_at": "2025-03-01",
"department": "HR",
"allowed_groups": ["all-staff"]
}
Metadata does three jobs: filtering (search HR documents only, latest version only), access control (a user only ever retrieves chunks their groups allow) and citation (show the document title, the section, the link).
Why split at all
- A long passage produces a diluted embedding. One vector has to stand for too many ideas, so it ends up close to no specific question.
- Context limits and cost. Five 400-token passages are cheaper and far more focused than five 20-page documents.
- Precise citation. You can point at a section rather than “somewhere in this document”.
But passages that are too short break in the other direction: “The deadline is 30 days.” on its own does not say the deadline for what.
Chunking strategies
| Strategy | How it works | Suits | Watch out for |
|---|---|---|---|
| Fixed size | Cut every N tokens, overlapping M | A starting point; uniform text | Cuts mid-sentence and mid-table |
| Recursive | Try paragraph → sentence → word until it fits | Most prose | A good default; what most libraries ship |
| Structural | Split on headings, sections, clauses | Policies, technical docs, Markdown/HTML | An over-long section still needs splitting |
| Semantic | Split where the embeddings of successive sentences diverge | Long text that keeps changing subject | More compute; measure before you believe it helps |
| Parent–child (small-to-big) | Retrieve on small chunks, put the larger parent in the prompt | Needing precise retrieval and full context at once | You must store the parent–child relation |
The trick worth knowing: put the context inside the chunk
Before embedding, prepend the document title and the section path to each chunk. “The deadline is 30 days.” becomes:
Expense reimbursement policy › 4. Deadline for submitting receipts
The deadline is 30 days from the date the expense was incurred.
One extra line, and both the retriever and the model now know what the passage is about.
Overlapping chunks, in code
def chunk_words(text, size=120, overlap=20):
words = text.split()
step = size - overlap
chunks = []
for start in range(0, len(words), step):
piece = words[start:start + size]
if piece:
chunks.append(" ".join(piece))
if start + size >= len(words):
break
return chunks
The overlap repeats the last few sentences of one chunk at the start of the next, so an idea sitting exactly on a boundary still appears whole in at least one chunk. Run on a 1,000-word document this returns 10 chunks: nine of 120 words and a last one of 100, with exactly 20 words shared between each neighbouring pair, and every word present in at least one chunk. (In production, count the tokens of your embedding model rather than whitespace-separated words.)
Choosing size and overlap
- A reasonable starting point: 200–500 tokens per chunk, overlapping by 10–20%.
- Fact-lookup questions (“what is the allowance?”) usually favour smaller chunks.
- Explanatory questions (“how does the approval process work?”) usually favour larger chunks, or parent–child.
- Never exceed the embedding model's input limit — the excess can be truncated silently.