Back to News & insightsGuides

Document chunking: preserve meaning before optimizing retrieval

Design retrieval units that keep headings, exceptions, tables, and source identity intact, then evaluate whether the returned evidence can actually answer the question.

Editorial guide · Updated September 28, 2026 · 8 min read
A silver book separates into curved page sections with interlocking edges.

A retrieval system can return the correct words and still provide the wrong evidence. A sentence may lose the heading that limits its scope. A table row may arrive without the column labels that explain its numbers. An exception may be separated from the rule it changes. These are often chunking problems before they are embedding or generation problems.

Chunking decides which pieces of a document become independently retrievable units. It is therefore an editorial and information-structure decision as well as a technical one. The aim is not to discover one universally optimal token count. It is to create units that remain meaningful when removed from their original surroundings and presented to a model answering a specific question.

Begin with the evidence a reader would need

Imagine a hypothetical support assistant for a collection of camera manuals. A user asks whether a certain recording mode works with an external microphone. The answer may depend on the camera revision, the recording mode, an adapter requirement, and a footnote in a compatibility table.

Before choosing a splitter, identify the smallest complete evidence package a knowledgeable support person would need. It might include a section heading, two paragraphs, a table row with headers, and the footnote. A fixed-length window that happens to capture only the row is not enough, even if its embedding is highly relevant.

Collect representative questions first. Include direct factual questions, comparisons across sections, and requests that require an exception. This gives the chunking experiment a purpose. Without those questions, teams often optimize convenient properties of the index while never checking whether the retrieved units can support a correct answer.

Structure can carry more meaning than proximity

Documents contain relationships that plain text extraction can flatten. A heading applies to the paragraphs beneath it. A caption explains an image. A footnote qualifies a nearby value. A list may contain alternatives rather than steps. Preserve those relationships where the source format allows it.

Store a hierarchy path with each chunk, such as manual, chapter, section, and subsection. Include document identity and version separately from the text. If the same sentence appears in two camera manuals, the source identity may determine which one answers the question. Similar wording does not make the evidence interchangeable.

Treat extraction quality as an upstream dependency. Broken reading order, missing symbols, and merged columns cannot reliably be repaired by adjusting chunk size. Inspect representative extracted documents before indexing the full collection. A small manual review can reveal errors that would otherwise become thousands of misleading retrieval units.

Contextualization is useful but must remain distinguishable

Anthropic's contextual-retrieval article describes adding a short contextual explanation to chunks before indexing, helping locate them within the broader document. This is one approach to information lost when passages are isolated. Its reported results belong to its stated evaluation, rather than establishing a universal improvement for every collection.

Anthropic: Engineering research and guidance

If you generate contextual descriptions, keep them separate from original source text. A model-written explanation can help retrieval but can also introduce an incorrect assumption. The final answer should not cite generated context as though it were a sentence written by the manual's author.

For the camera assistant, a safe context field might identify the relevant model and chapter using verified metadata. A generated claim that the passage applies to all recording modes would require evidence. Context should clarify the source boundary, not silently broaden it.

Compare splitting policies under a fixed evidence budget

Try a small number of clear policies: bounded paragraph groups, sections with a size limit, or smaller retrieval units that expand to their parent section when selected. Keep the retrieval model and evaluation questions stable while comparing them. Changing every component simultaneously makes it difficult to learn what helped.

Use the same final context budget. Returning more text can make one policy look better simply because it supplies substantially more evidence. Measure both answer support and the amount of material delivered to the generator. A useful chunking policy should use that limited space efficiently.

Record how often important evidence is split across units. Some questions legitimately need several chunks, but repeated fragmentation can consume the retrieval budget with overlapping pieces. Inspect the actual returned set, not just whether one relevant chunk appeared somewhere among many candidates.

Overlap can rescue a boundary and create another problem

Overlapping windows help when a relevant sentence sits near a cut. However, excessive overlap can fill the result set with nearly identical passages. The generator then sees repeated evidence while a necessary exception or second source is pushed out of the context budget.

Deduplicate returned passages before assembling the final context, while preserving source locations. If two overlapping chunks are merged, verify that the combined text remains in the correct order. Avoid stitching separate sections into a paragraph that the original author never wrote.

Measure overlap as part of the complete retrieval workflow. An index with more chunks increases storage and ingestion work, and a reranker may spend effort scoring redundant candidates. Those costs can be justified, but only if the resulting evidence improves the task rather than merely increasing the number of searchable fragments.

Tables deserve a specific representation

A table is not just a collection of short paragraphs. A value depends on its row label, column label, units, and sometimes a spanning header or footnote. Extracting each cell independently can destroy the information needed to interpret it correctly.

For the camera manual, represent a compatibility row with its headers and relevant notes. Preserve the original table location so a reader can verify the answer. If the table is too complex to flatten reliably, retain an image or structured representation and route the question through a workflow designed to inspect it.

Test with deliberately similar rows. Two recording modes may differ by one symbol or one footnote marker. A system that answers easy rows correctly can still fail precisely where the table was designed to communicate an exception. Include those distinctions in the evaluation set from the beginning.

Versions and permissions travel with every chunk

A chunk should retain the identity, version, and access conditions of its source. Reindexing must not create an orphaned fragment that remains searchable after the document is withdrawn or restricted. Retrieval quality includes returning evidence the user is actually allowed to see.

When a manual changes, decide whether old versions remain available and how questions identify the intended one. A user asking about an older camera revision may need the older manual. Simply deleting every previous version can be as incorrect as mixing all versions without labels.

Keep deletion and replacement operations testable. Select a source document, update or withdraw it, and verify which chunks remain in the index and caches. The pipeline should be able to explain the result without relying on a manual search for fragments with similar text.

Evaluate support before evaluating prose

For each test question, identify the source material needed for a correct answer. Then inspect the retrieved context before generation. Does it contain the relevant rule, the applicable exception, and enough identity information to choose the right version? If not, a fluent final answer cannot repair the missing evidence reliably.

Separate retrieval failures from answer-composition failures. The correct evidence may be present while the generator ignores a footnote. Conversely, the generator may reason carefully over incomplete context and still reach the wrong result. These failures require different fixes, so preserve the intermediate artifacts during evaluation.

Include unanswerable questions. A manual may not specify compatibility with a third-party accessory. The appropriate outcome is a clear statement of missing evidence, not a guess assembled from vaguely related chunks. A retrieval system should help establish when the collection does not support an answer.

Inspect boundary cases with a small review worksheet

For every failed question, record the source section, the chunk boundaries, the returned candidates, and the final context. Mark the first point where necessary meaning was lost. This produces a concrete repair target, such as retaining table headers or expanding a selected paragraph to include its exception.

Prioritize fixes that apply to a meaningful class of documents. A one-off rule for one paragraph may improve the test set while making the pipeline harder to maintain. Prefer a documented structural rule, then test it on other manuals with similar layouts and on documents where the rule should have no effect.

Finally, rerun the same questions after the change and inspect new failures. Larger chunks may restore one exception while crowding out another source. Chunking decisions involve tradeoffs, and the review process should make those tradeoffs visible rather than claiming one size is optimal everywhere.

Keep the index connected to the original document

The best retrieval unit is still a representation of a source, not a replacement for it. Provide a reliable path from the answer to the original passage, table, or page. Readers should be able to verify a claim in context, especially when the answer depends on a subtle qualification.

Choose a chunking policy that fits the document types, question patterns, and operating constraints of the application. Revisit it when those change. Adding slide decks, scanned manuals, or multilingual documents may introduce structural issues that the original policy never encountered.

Good chunking preserves enough meaning that retrieval can do its job honestly. It makes relevant evidence discoverable, keeps qualifications attached, and helps the generator recognize what the source can and cannot establish. That foundation often matters more than another round of prompt polishing after the evidence has already been damaged.

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.