Chunking Strategies for Enhanced RAG Quality
A lot of RAG systems struggle not because the model is weak, but because the source data was split badly. Chunk too aggressively and you lose context; chunk too broadly and retrieval turns noisy. Here's a technical map of ten chunking strategies — what each one does, and how, when and where to use it.
Chunking is usually treated as a default setting — pick a chunk size, add a bit of overlap, move on to the embedding model. That's a mistake. The chunk boundary decides what your retriever can actually find, what context reaches the LLM, and ultimately how grounded the final answer can be. Two systems using the identical embedding model and the identical LLM can produce very different answer quality purely because one of them chunked its source documents better.
There is no universal "best" chunk size. The right choice depends on document structure, query patterns, context requirements, latency, cost, retrieval method, and — ideally — actual evaluation results rather than a guess. This piece walks through ten chunking strategies grouped into three families, gives a decision framework for picking one, and closes with what controlled research has actually found when these strategies are measured head-to-head.
01Ten Strategies, Three Families
Structural strategies split on visible or countable boundaries. Meaning-aware strategies split — or repair — based on what the text actually means. Retrieval-architecture strategies change what gets indexed versus what gets returned.
Structural & Rule-Based Chunking
1. Fixed-Size Chunking
Use for: simple, uniform textSplits text into chunks of a fixed number of characters or tokens, usually with a fixed overlap. Simple, fast and fully deterministic, with no dependency on document structure — but naive fixed-size splitting with large chunks and heavy overlap is exactly what underperformed in controlled testing (see Section 3). Tuned carefully — smaller chunks, little or no overlap — it stays a reasonable baseline; used as a lazy default, it usually isn't.
2. Recursive Chunking
Use for: general-purpose documentsRecursively splits on a prioritized list of separators — paragraph breaks first, then sentences, then words — until each chunk fits a target size, preserving more natural boundaries than raw character counting. This is what LangChain's RecursiveCharacterTextSplitter implements, and it's the de facto default for mixed or unknown document types because it degrades gracefully without needing document-specific logic.
3. Sentence-Based Chunking
Use for: FAQs and support contentSplits at sentence boundaries using NLP sentence segmentation, so each chunk is a self-contained thought. Well matched to FAQ pages and support articles, where each retrieval unit should map cleanly onto one question's answer. The trade-off is that very short chunks can lack surrounding context — a problem sliding-window or contextual chunking (below) exist to fix.
4. Paragraph-Based Chunking
Use for: reports and manualsSplits at paragraph boundaries, preserving the natural unit of argument or explanation a human author already delineated. A good fit for reports, manuals and long-form prose where paragraphs are internally coherent — though resulting chunk sizes are less uniform and less predictable than fixed-size or recursive splitting, which complicates capacity planning for the vector index.
5. Sliding Window Chunking
Use for: content whose context crosses boundariesChunks with deliberate overlap — a window that advances a fixed number of tokens at a time — so content near a boundary still appears in full within an adjacent chunk. Useful whenever meaning-carrying context crosses a natural split point, such as a clause continuing an argument from the previous paragraph, at the cost of a larger index and some retrieved redundancy.
6. Document-Structure-Aware Chunking
Use for: PDFs, policies, and technical docsUses the document's own markup — headings, sections, list items, table boundaries — as chunk boundaries instead of raw character counts. Close to essential for PDFs, policies and technical documentation, where a chunk that silently crosses a section heading mixes two unrelated topics into one embedding. Document-parsing pipelines such as Unstructured's partitioning tools exist specifically to recover this structure from PDFs and HTML before chunking begins.
Meaning-Aware Chunking
7. Semantic Chunking
Use for: topic-heavy contentInstead of splitting on visible structure, semantic chunking embeds consecutive sentences and inserts a boundary wherever the similarity between neighboring sentence embeddings drops — wherever the topic actually shifts, regardless of paragraph or heading markers. The approach traces back to Greg Kamradt's widely cited "five levels of text splitting" work; a cluster-based variant that maximizes embedding similarity within each chunk globally, rather than only comparing adjacent sentences, achieved the highest precision in Chroma's controlled chunking evaluation (Section 3). Best for long-form articles, transcripts and meeting notes that don't map cleanly onto headings.
8. Contextual Chunking
Use for: chunks whose meaning depends on surrounding textPrepends a short, LLM-generated explanation of what a chunk is about and where it sits in the source document, before that chunk is embedded or indexed — so a fragment reading "revenue grew 3% over the previous quarter" gets rewritten with the company name and fiscal period attached. This is Anthropic's Contextual Retrieval technique: contextual embeddings alone cut retrieval failures by 35%, adding contextual BM25 lexical search on top cut failures by 49%, and layering in reranking pushed the improvement to 67% (see Section 3). Reach for it when chunks are frequently ambiguous out of context — financial filings, legal contracts, dense technical specifications.
Retrieval-Architecture Chunking
9. Parent-Child Chunking
Use for: precise retrieval with broader contextIndexes small, precise chunks for matching, but returns the larger "parent" chunk or section they belong to as the context actually sent to the LLM. LangChain's ParentDocumentRetriever and LlamaIndex's small-to-big retrieval pattern implement this directly: small chunks give the retriever precision — a query matches one specific sentence — while the larger parent gives the generation step the surrounding context it needs to actually answer well. It decouples what gets matched from what gets read.
10. Hierarchical Chunking
Use for: large repositoriesExtends parent-child chunking into several nested levels — document, section, paragraph, sentence — letting retrieval operate at whichever level best matches a given query and letting the system walk up or down the hierarchy as needed. Suited to large repositories — documentation sites, knowledge bases spanning thousands of documents — where one flat chunk size can't serve both broad "which document covers this" queries and narrow "what exactly does this clause say" queries equally well.
02A Decision Framework
Matching content type to a starting strategy — treat this as a starting point for evaluation, not a final answer.
| Content Type | Start With | Why |
|---|---|---|
| Logs, transcripts, uniform text | Fixed-size or recursive | No exploitable structure; predictable chunk size matters more than boundary precision |
| FAQs, support articles, Q&A pairs | Sentence-based | Each retrieval unit should map to one self-contained answer |
| Reports, manuals, long-form prose | Paragraph-based or recursive | Paragraphs are already coherent argument units |
| PDFs, policies, technical specifications | Document-structure-aware | Section and heading boundaries carry meaning fixed-size splitting would destroy |
| Long-form articles, meeting notes without headings | Semantic chunking | Topic shifts don't align with any visible structural marker |
| Financial filings, legal contracts, dense specs | Contextual chunking | Individual chunks are frequently ambiguous without surrounding context |
| Precision-sensitive QA over long documents | Parent-child chunking | Retrieval precision and generation context are different requirements |
| Large multi-document knowledge bases | Hierarchical chunking | Query granularity varies too widely for one flat chunk size |
| Content where key context sits past a boundary | Sliding window | Deliberate overlap prevents boundary-adjacent meaning from being cut |
| General-purpose default / mixed document types | Recursive | Degrades gracefully without needing document-specific logic |
03What the Research Actually Shows
Two controlled studies are worth knowing well, because their numbers are specific enough to act on rather than just directionally suggestive.
Chroma's Chunking Evaluation
Chroma's technical report on chunking built a token-level evaluation framework — precision, recall and intersection-over-union against known-relevant spans — across 472 queries over five different corpora (state documents, Wikipedia, finance, biomedical literature and chat logs). The headline finding: chunking strategy choice changed recall by as much as 9 percentage points between methods. A cluster-based semantic chunker capped at 200 tokens produced the highest precision of everything tested; an LLM-driven chunker that used a language model to choose split points directly produced the highest recall. Recursive character splitting at a modest 200-token size with no overlap held up well against far more sophisticated methods. The most notable result may be the negative one: OpenAI's own commonly used default settings — 800-token chunks with 400-token overlap — scored the lowest across nearly every metric in the study, a reminder that a popular default is not the same thing as an evaluated one.
Anthropic's Contextual Retrieval
Anthropic's Contextual Retrieval work targeted a specific failure mode directly: a chunk that is perfectly well-formed but meaningless once separated from its source, because the company name, date, or subject it refers to lived in a different chunk. The fix — an LLM-generated explanatory prefix added to each chunk before embedding and indexing — produced large, layered improvements.
Both studies point the same direction: the marginal chunking decision —200 tokens versus 800, semantic versus fixed, contextualized versus raw — is not a minor tuning knob. It is frequently a bigger lever on end-to-end RAG quality than swapping the embedding model or the LLM itself.
04Test It, Don't Assume It
The practical takeaway from both studies isn't "always use strategy X" — it's that chunking strategy needs the same evaluation discipline as any other model choice. A minimal version of that discipline: build a small labeled set of queries with known-relevant source spans, measure precision and recall (or IoU) for each candidate chunking strategy against that set, and only then pick a default for production. That evaluation loop is cheap relative to the cost of silently shipping a chunking strategy that caps retrieval quality no downstream reranking or prompt engineering can fully undo.
If you are building a RAG system, chunking is one layer worth testing deliberately — document structure, query patterns, context requirements, latency, cost and retrieval method all pull in different directions, and the strategy that wins for one corpus is rarely the one that wins for the next.
05Further Reading
The primary sources behind the research cited above.