Azure AI Search index projections shape enriched content into one-to-many search documents, which is especially important for RAG systems that split a parent document into chunks. The skillset defines projection selectors that map chunk-level content and repeated parent metadata into a target index. Microsoft currently recommends a single chunk-oriented index with parent fields repeated on every chunk and projectionMode set to skipIndexingParentDocuments for most classic RAG scenarios.
Within Microsoft AI Agents, index projections are the bridge between ingestion and retrieval. They decide whether every searchable chunk carries the title, source, tenant, ACL, document ID, or other parent metadata an agent needs later.
Enterprise RAG chunking provides the design context; projections make the resulting chunk model concrete in Azure AI Search.
Projection design starts from the query document shape
Before writing a skillset, decide what one search result should look like. In a typical chunk index, each result needs a unique chunk key, parent ID, chunk text, vector, and selected parent metadata.
The index should be designed for retrieval first; the projection then maps enrichment-tree values into that schema.
A projection is not a join feature. Query-time joins are not available in classic Azure AI Search, so required parent context should be present in the searchable document.
A parent key keeps chunks traceable to the source document
The chunk index needs a field associating every child with its parent, commonly called parent_id.
Microsoft guidance requires this association field to be a string and filterable, but it should not be the document key itself.
This enables grouping, filtering, deletion workflows, and reconstruction of which source document produced a retrieved chunk.
Repeating parent fields is usually simpler than mixed shapes
When title, source URL, tenant, language, or security metadata repeats on every chunk, all documents in the index have a uniform shape.
This costs some storage but makes queries, filters, result rendering, and permission checks much simpler.
For most RAG workloads, Microsoft recommends this pattern over adding separate parent documents to the same index.
skipIndexingParentDocuments avoids null-heavy parent records
The default projection behavior can index parent documents in addition to children, creating search documents where chunk fields are null.
If five source documents produce one hundred chunks, that can produce 105 index documents with two different shapes.
Using skipIndexingParentDocuments keeps the index chunk-centric and avoids accidental retrieval of parent-only rows.
Separate parent and child indexes add query complexity
Azure AI Search can project child chunks to one index while an indexer populates a separate parent index.
This can be useful when parent metadata has a distinct lookup lifecycle, but classic search has no query-time join across indexes.
Applications must perform their own second lookup or duplicate needed metadata anyway, so use two indexes only when that separation has clear value.
sourceContext defines the projection granularity
A selector’s sourceContext points to the enrichment-tree path that represents each child unit, for example a page or chunk array.
Mappings under the selector pull values relative to that context into target fields.
If source context is wrong, one parent can collapse into one document or fields can repeat at the wrong level, so inspect enrichment outputs during development.
Projection mappings should include security metadata
Permission-aware RAG needs ACL or tenant fields on every projected chunk, not only on the parent.
If chunk retrieval returns text without the parent’s access labels, the application has to reconstruct security context after retrieval, which is easy to get wrong.
Azure AI Search Filtered Vector Search explains why these projected fields must be filterable at query time.
Integrated vectorization fits naturally into projections
A skillset can split content, generate embeddings, and project both text and vector fields into the chunk index.
The embedding output should map to a vector field whose dimensions and profile match the model used.
Keep the nonvector chunk text too; vector search returns document fields, not a reverse conversion from the embedding.
Changes and deletions should preserve parent-child consistency
Indexer change tracking can update projected chunks when a parent document changes, but deletion behavior depends on the source connector and key strategy.
Use stable parent IDs and source deletion tracking where supported, and test what happens when a parent shrinks from ten chunks to six.
Orphaned chunks can create stale retrieval even when the latest parent content looks correct.
Reset/reindex carefully after projection-schema changes
Changing chunking, mappings, projection mode, or index schema often requires a reset and rerun so old documents are rebuilt consistently.
A reset clears indexing state but does not automatically remove every orphaned document in all scenarios, so validate the target index document set after structural changes.
Azure AI Search Indexer Failures covers reset/run and troubleshooting behavior.
Index projections are successful when every retrieved chunk is independently useful and governable
The mature projection contains chunk text/vector, stable key, parent ID, title/source, permissions, and any metadata needed to rank, filter, display, or trace the result.
Good projection design turns a complex enrichment tree into a search document shape an agent can trust without extra joins or guesswork.
Projection design should keep one stable document key per chunk. A key derived from a parent ID plus deterministic chunk identity helps updates overwrite the intended search document instead of creating duplicates after every rerun. Random chunk IDs are convenient during prototyping but make reconciliation and deletion harder when chunk boundaries change.
Chunk identity is especially important when the chunking algorithm evolves. Changing page size, overlap, tokenization, or document parser can produce a different number of child documents. A blue-green index rebuild is often cleaner than trying to mutate old child records in place because the new projection can be validated independently before query traffic switches.
Parent metadata should be curated, not copied wholesale. Repeating every parent field on every chunk can increase index size and update cost significantly. Include fields needed for security, filtering, ranking, result rendering, traceability, or downstream navigation. Keep large binary or verbose metadata in the source system and store only the stable pointer required to retrieve it later.
Search result rendering should use projected source metadata rather than inventing context from chunk text. Title, page, section, document URL, and parent identifier let the agent or UI cite the source accurately. This becomes particularly important when several chunks contain similar boilerplate and the user needs to understand which original document supports the answer.
Projection mappings should be tested with documents that create zero, one, and many chunks. Empty documents, images without extracted text, very short files, and parser failures can create enrichment paths that do not match the selector’s assumptions. Inspect execution warnings and sample the resulting index instead of assuming every parent produces the same tree shape.
When integrated vectorization is used, the text chunk and its vector should be generated from the same logical content. If preprocessing sends one normalized string to the embedding skill but projects a different string to the searchable chunk field, retrieved text can appear unrelated to the vector that caused the match. Keep transformations explicit and versioned.
Projection changes should have regression queries. Before rebuilding production, run a fixed set of queries against old and new indexes and compare answer-supporting chunks, metadata completeness, filter behavior, and latency. A technically correct projection can still reduce retrieval quality if chunk boundaries or parent fields change in ways the ranking pipeline does not expect.
For multi-tenant systems, projected parent IDs and chunk keys should avoid collisions across tenants. Prefix or otherwise namespace identifiers so two customers can never create the same search key accidentally. Combined with tenant filters, stable namespacing makes both deletion and incident investigation much safer.
Index projections should be included in data-contract testing. Given one known parent document, the pipeline should produce an expected number of child documents with the correct parent ID, security fields, source metadata, chunk text, and vector dimensions. This catches mapping drift before a production index is silently populated with malformed chunks.
When parent metadata changes more often than content, repeated fields mean many child documents can require updates. That storage/update amplification is usually acceptable for RAG simplicity, but it should be understood for high-churn sources. If metadata changes constantly at massive scale, evaluate whether a different index design or application-side lookup provides a better operational trade-off.
Projection definitions are part of the search application’s source code. Version the skillset, index schema, and indexer together, and keep a migration note explaining which query code expects which shape. A deployment that updates the skillset without the matching index/query changes can produce data that is technically indexed but unusable by the application.
Projection migrations should also include deletion testing. Replace or delete a source parent, rerun the indexer, and verify old child chunks no longer remain searchable under stale metadata. If the source connector cannot guarantee that behavior, add explicit cleanup or rebuild procedures. RAG correctness depends as much on removing obsolete chunks as on adding new ones.
Keep projection tests in CI alongside the skillset and index schema so a field rename or chunk-path change cannot reach production without proving the expected child documents still materialize correctly.