Chunking is not a preprocessing detail; it defines the units a retriever is allowed to return. The live Databricks Generative AI Engineer Associate exam guide now explicitly calls for choosing chunking strategies based on document structure, model constraints, and retrieval evaluation, including advanced chunking approaches. That framing is useful because there is no universally correct token size.
Chunk quality depends on what users ask. A policy question may need one coherent section, a troubleshooting question may need a short procedure, and a table lookup may need row/column context that a naive text splitter destroys.
The right method is experimental and data-aware: understand the source structure, choose candidate chunk boundaries, preserve metadata, build the index, test retrieval on representative questions, and adjust from evidence. The general principle behind data-quality accountability applies because bad chunking is a data-quality defect at the retrieval layer.
Start from the answerable unit
Ask what smallest source unit can answer one representative question without requiring missing context from neighboring text.
That unit might be a paragraph, a subsection, a table row plus headers, a code function, or several connected sentences.
Chunk size should follow the semantic unit first and token limits second where the document structure makes that possible.
The answerable unit can also depend on how evidence is cited. If the application must link users back to a page or section, chunks should preserve source boundaries that make the citation meaningful. A semantically perfect chunk assembled from fragments across distant pages can be difficult for a reviewer to verify.
Answerability should be tested on multi-part questions too. A user may ask for two facts stored in different sections. One large chunk can contain both but retrieve poorly; two smaller chunks can retrieve independently and require the model to combine evidence. Evaluation should include questions whose evidence naturally spans several chunks.
Fixed token windows are a baseline, not a doctrine
Fixed-size windows are easy to implement, predictable, and useful as a starting benchmark.
They can split headings from content, cut lists in half, separate table headers from values, or combine unrelated sections when the token boundary ignores structure.
Use them when sources are regular or when a baseline is needed, then compare against structure-aware alternatives rather than assuming simplicity means optimal retrieval.
Fixed windows are still useful as a control group. Before investing in complex parsers, measure how a simple baseline performs. An advanced strategy that is harder to maintain should produce a meaningful retrieval improvement; otherwise the simplicity of fixed windows can be an operational advantage.
Window baselines should be tested at more than one size. Comparing 256, 512, and 1,024-token windows can show whether the corpus is sensitive to context width before engineering a custom parser. A simple curve of retrieval quality versus record count gives a useful reference for judging more complex strategies.
Overlap protects context and creates duplication
Overlap can preserve meaning when an important sentence sits near a chunk boundary, but every repeated token increases record count, storage, embedding work, and duplicate retrieval.
Large overlap can cause the top-k result set to contain several near-identical chunks, reducing evidence diversity.
Measure whether overlap improves retrieval recall enough to justify the extra index size and repeated context.
Overlap can interact with reranking and top-k limits. If three nearly identical overlapping chunks occupy the candidate set, a reranker has less opportunity to surface a different supporting passage. Deduplication or diversity-aware selection may be needed when overlap is intentionally large.
Structure-aware chunking preserves document semantics
Use headings, Markdown structure, HTML sections, PDF layout, code blocks, tables, or other source-specific boundaries where available.
Carry parent section title, document ID, page, timestamp, access classification, and business metadata into the chunk record.
Metadata supports filtering, citation, governance, and debugging and often improves retrieval quality more safely than simply enlarging every chunk.
Document structure should be normalized across versions. If one author uses proper headings and another exports flat text, structure-aware chunking can produce inconsistent units for similar documents. Preparation pipelines should detect weak structure and fall back to another strategy rather than trusting malformed markup blindly.
Structure-aware parsing also needs a fallback for corrupted or inconsistent documents. If headings are missing or OCR merges sections incorrectly, a hierarchy-based splitter can create enormous or meaningless chunks. Detect abnormal length and structure distributions and route those documents through a safer secondary method.
Chunk size interacts with embedding-model context
Do not create chunks larger than the embedding model can represent properly. Truncation can silently remove the part containing the answer.
Even below the model limit, larger chunks can produce embeddings that average several topics into one representation.
Choose the smallest chunk that preserves the necessary semantic context, then use parent-child or post-retrieval context expansion when the application needs broader reading context.
Embedding context limits should be checked after preprocessing, not from raw character counts. Tokenization varies by language, code, punctuation, and model. Measure actual token length distributions so rare oversized chunks do not silently truncate while average chunks appear safe.
Retrieval evaluation should drive the decision
Create a set of realistic questions with known relevant source passages or subject-matter-expert judgments.
Compare candidate chunking approaches using retrieval metrics and qualitative review. Inspect failures: missing evidence, fragmented evidence, duplicate results, or irrelevant broad chunks.
One global average can hide categories. Legal policies, tables, support procedures, and free-form notes may deserve different strategies in the same application.
Evaluation should include downstream answer quality as a secondary measure. Two chunking strategies can have similar retrieval recall while one produces cleaner context that the model uses more effectively. Retrieval-first metrics locate search quality, and end-to-end metrics reveal whether that quality actually improves answers.
Evaluation should inspect the retrieved text, not only metric scores. A high recall result can still return fragmented or confusing evidence that causes generation errors. Human review of representative candidate sets often reveals duplication, missing headers, and table context problems that a numeric score compresses away.
Reranking can rescue candidates but not missing evidence
Rerankers can reorder an initially retrieved candidate set using a more expensive relevance model.
They are useful when vector similarity finds broadly relevant content but ranks the strongest evidence too low.
Reranking cannot recover a passage that chunking and first-stage retrieval never placed in the candidate set. Fix source and chunk design before using reranking as a universal patch.
Reranking candidate size is another lever. A larger first-stage candidate set can increase the chance that the correct passage reaches the reranker and adds latency and cost. Tune candidate count and final top-k together rather than changing the reranker model in isolation.
Cost appears in indexing and generation
More and smaller chunks increase embedding volume, index size, sync work, and potential query candidates. Larger chunks can reduce record count and increase prompt tokens when retrieved.
The same scale thinking behind big-data analytics applies: data shape changes compute and storage cost as the corpus grows.
Measure end-to-end cost per useful answer rather than optimizing the number of chunks in isolation.
Scale tests should estimate record growth from overlap and average chunk size before indexing a full corpus. A small prototype can hide a tenfold expansion that later affects index capacity, sync duration, and cost. Record-count modeling is part of chunking architecture when the source estate is large.
Estimate the operational cost of index rebuilds when changing chunking. A strategy that creates twice as many records affects embedding time, sync duration, vector storage, backup/recovery, and evaluation runtime. Those costs are acceptable when quality improves enough, but they should be visible before the full corpus is regenerated.
Version chunking as part of the retrieval product
Record the splitter version, parameters, source version, embedding model, and index version so experiments can be reproduced.
Rebuild or maintain parallel indexes when changing chunking materially, then route controlled evaluation traffic and compare results.
Chunking improves production retrieval when teams can explain why a boundary exists, prove that it retrieves evidence more reliably, and roll back when the new strategy creates worse behavior.
Version transitions need a cleanup plan. Parallel indexes are useful for evaluation, but old indexes consume cost and can be queried accidentally if clients retain stale configuration. After rollback windows close, deprecate aliases, remove obsolete endpoints, and archive the experiment metadata needed to reproduce the decision.
Production migration should also update caches and references keyed to old chunk IDs. Citation links, feedback stores, saved examples, and evaluation labels can all point at the previous segmentation. A chunking change is therefore a data-contract change for every system that refers to retrieved units by identifier.
Chunking should also preserve tables and list semantics deliberately. A table split into independent rows can lose the headers that give values meaning; a numbered procedure split mid-sequence can make step four look actionable without steps one through three. Specialized handlers are justified when document structure carries meaning that plain text windows erase.
Multilingual documents create another boundary problem because token density varies by language. A fixed token limit may produce very different semantic size across English, Japanese, code, or mixed content. Evaluate chunk statistics and retrieval separately by language when the corpus is multilingual.
Finally, chunking rules should be easy to explain to content owners. When writers understand that headings, concise sections, stable tables, and metadata improve retrieval, source authoring itself can become more retrieval-friendly and reduce the amount of downstream parsing needed.
One more useful check is source-author behavior over time. If authors change templates, export formats, or table layouts, chunking assumptions can degrade even though the splitter code is unchanged. Monitor chunk-length and structure distributions so source-format drift becomes visible before retrieval quality collapses.