A production RAG application on Databricks is not one retriever and one model call. It is a chain of governed data, extraction, chunking, embeddings, an AI Search index, retrieval logic, optional reranking, prompt assembly, model or agent execution, serving, evaluation, monitoring, and access control. The live March 2026 Databricks Generative AI Engineer Associate exam guide explicitly spans those layers, which makes architecture reasoning more valuable than memorizing isolated product names.
The broader big-data analytics perspective is useful because RAG quality begins with data engineering. Documents arrive from business systems, are cleaned and normalized, stored in governed tables or volumes, transformed into chunks, and indexed with metadata. Retrieval can only return what that preparation made visible.
The design review should therefore follow one user question end to end and ask where quality, latency, cost, security, and freshness can change. That path reveals dependencies that a clean RAG diagram usually hides.
Source quality defines the ceiling
Select sources that actually contain the knowledge the application needs and that can be used legally and securely.
Remove boilerplate, duplicated pages, navigation artifacts, corrupted scans, or irrelevant content that lowers retrieval precision.
Keep source identity and update timestamps so operators can trace an answer back to the current document and detect when the corpus is stale.
Source selection should include authority ranking. Internal policy, product documentation, support tickets, and user-generated notes can conflict. Retrieval needs a rule for which source wins when two chunks disagree, and that authority should appear in metadata or ranking logic rather than being left to the model to infer from writing style.
Source authority can also affect retrieval ranking directly. A current approved policy should outrank an old draft even when the draft has wording closer to the user’s question. Metadata for approval status, effective date, and document class gives the retrieval layer signals that semantic similarity alone cannot infer reliably.
Governed preparation turns documents into retrievable units
Extraction and parsing should preserve useful structure such as headings, tables, page references, and metadata where the use case needs them.
Store normalized content in governed Delta tables or volumes with permissions and lineage. The pipeline should be reproducible so a parsing change can rebuild the corpus consistently.
Python remains a practical transformation layer, and the broader role of Python in data science applies: code should make preprocessing assumptions explicit rather than hiding them in manual notebook steps.
Extraction pipelines should keep raw and normalized representations when auditability matters. If parsing loses a table column or changes reading order in a PDF, operators need the original to diagnose the defect and improve the parser. Storing only cleaned text makes quality issues harder to trace because the evidence that was discarded is no longer available.
Chunking encodes a retrieval hypothesis
Chunk size and overlap decide how much semantic context each indexed unit represents. Too small can fragment meaning; too large can dilute the matching signal and waste model context.
Chunk around document structure when possible—sections, paragraphs, tables, or semantic boundaries—rather than using one arbitrary token count for every source.
Evaluate chunking with realistic questions. A chunking strategy is correct only if it helps retrieval recover the evidence users need.
A production corpus can use multiple chunking strategies at once. Tables, long narrative sections, FAQs, and code examples may need different boundaries and metadata. The index can still present one retrieval interface while preparation logic chooses source-specific methods. The important operational requirement is that the strategy used for each chunk remains identifiable during debugging.
Chunk identifiers should remain stable across harmless rebuilds when possible. Stable IDs make feedback, citations, cached references, and evaluation datasets easier to reconcile. When chunk boundaries change materially, create a new version namespace so old evaluation results are not mistakenly compared against different retrieval units.
Embeddings create the comparison space
An embedding model converts text into vectors whose geometry represents semantic similarity according to the model’s training and architecture.
Model context length, embedding dimension, language coverage, domain behavior, latency, and cost all affect the retrieval system.
Version the embedding model with the index. Changing embeddings without rebuilding or clearly separating indexes can create inconsistent similarity behavior.
Embedding migration should be planned like a schema migration. Build the new vectors in parallel, create a separate index, run retrieval evaluation, and switch traffic only after evidence shows the new representation is better. Replacing vectors in place without an evaluation window makes rollback difficult and can mix representations unexpectedly.
AI Search is the retrieval service, not the whole RAG system
Databricks AI Search, formerly Vector Search, can index Delta-backed data and provide vector, hybrid, filtering, full-text, and reranking capabilities depending on configuration.
Endpoint SKU, index size, dimensionality, query mode, requested result count, update frequency, and traffic all influence performance and cost.
Treat the index as a production dependency with ownership, sync health, capacity planning, access control, and a rebuild strategy.
AI Search access control should match corpus governance. If one index contains multiple tenants or sensitivity levels, filtering and permissions must prevent a user from retrieving unauthorized chunks before generation. Post-generation redaction is weaker because the model has already seen the restricted content.
Index security should be tested with adversarial queries. Attempt to retrieve another tenant’s document, content outside the caller’s classification, and documents that filters should exclude. A successful normal query proves relevance; negative tests prove isolation.
Retrieval quality needs direct measurement
Measure whether the right evidence appears in the candidate set before blaming the model for a bad answer.
Use recall- or ranking-oriented evaluation, annotated examples, human review, or task-specific metrics to compare chunking, embeddings, filters, hybrid search, and reranking.
Separate retrieval failure from generation failure. A perfect model cannot cite evidence it never received.
Evaluation sets should include ‘no answer’ cases where the corpus does not contain the requested fact. A retriever that always returns something can look productive while feeding the model irrelevant evidence. Measuring when the system should abstain is part of retrieval quality, especially in regulated or customer-facing applications.
Retrieval evaluation should include freshness cases. Ask questions whose answer changed recently and verify the newest authoritative passage is retrieved above obsolete text. This catches pipelines that index new documents while leaving older contradictory versions equally competitive.
Prompt assembly should make provenance visible
Retrieved chunks need clear delimiters, metadata, source references, and instructions about how the model should use them.
Control context size and remove duplicate or contradictory evidence where possible. The retrieval layer should deliver enough evidence without turning the prompt into an unfiltered document dump.
Where the application requires citations, preserve identifiers that can be returned to the user and verified.
Prompt assembly should include a policy for conflicting chunks. The model may receive two current-looking answers from different versions of a document. Ranking by authority, effective date, or source system can be safer than asking the model to reconcile silently. Conflict detection can also trigger escalation when the corpus itself needs cleanup.
Serving and access control are architecture decisions
Deploy the application or agent through a managed serving path with governed credentials and resource permissions.
The principles behind role-based access control matter because retrieval can expose documents the user should not see even when the model endpoint itself is secure.
Design authorization at source, index, tool, and application layers so one broad service credential does not turn all governed data into one shared prompt corpus.
Serving architecture should protect credentials used by the agent or application backend. Browser clients should authenticate to the application while the backend uses governed service identities or user context to access the serving endpoint and data. Exposing personal access tokens or broad service tokens in client code destroys the governance boundary.
Serving can also separate application identity from end-user identity. The backend might authenticate to Databricks with its own managed credential while carrying trusted user claims that drive authorization filters. Document which identity each layer uses so a broad backend permission does not erase per-user data policy.
Production RAG is a feedback system
Trace requests, log retrieval decisions, evaluate sampled outputs, monitor latency and cost, and incorporate subject-matter-expert feedback into the next iteration.
Use logging and monitoring as the operational baseline, then add RAG-specific evidence such as retrieved chunks, ranking, prompt version, model version, tool calls, and evaluation scores.
A production RAG stack is successful when teams can reproduce why an answer was generated, identify which stage failed, and improve that stage without guessing.
Feedback should be tied back to traceable components. A user saying ‘wrong answer’ is useful only when the system can recover the query, retrieved chunks, prompt, model, tools, and versions. That linkage turns anecdotal feedback into a reproducible test case that can improve data, retrieval, prompting, or generation separately.
Feedback triage should classify whether the complaint is source, extraction, chunking, embeddings, search configuration, prompt, model, or policy. That taxonomy turns production feedback into a queue for the correct owner and prevents every wrong answer from becoming a prompt-engineering task.
RAG architecture should also define deletion propagation. When a document is withdrawn, expires, or loses authorization, the corresponding chunks and vectors must disappear from retrieval within a known window. Retention and index sync are therefore part of security and governance, not only freshness.
Operational ownership should cover the full chain. One team may own source ingestion, another AI Search, another agent code, and another business policy. A production incident moves faster when each boundary has a named owner and a shared trace identifier instead of a generic ‘AI team’ queue.