RAG Evaluation: Retrieval and Answer Quality

A retrieval-augmented generation system can return polished prose while retrieving the wrong evidence. Good evaluation separates retrieval quality from generation quality, then measures citation support, safety, latency, and cost rather than collapsing them into one answer score. In the AWS-centered AIP-C01 objectives, RAG evaluation belongs with model/application testing and Amazon Bedrock Knowledge Bases. In Microsoft’s AI-103, the relevant work involves Microsoft Foundry evaluators, groundedness, retrieval relevance, and tracing deployed applications. Those are related but different implementation paths: an answer can be grounded in an incorrect document, and excellent retrieval can still lead to a fabricated response. The test design must identify which stage failed.

That pipeline view also explains why a RAG system can improve even when the foundation model never changes. Better chunk boundaries, cleaner metadata, stronger retrieval filters, a more appropriate embedding model, or a better reranking stage can all raise end-to-end quality. In Agentic AI Engineering, the diagnostic question is therefore not “Which model is smartest?” but “Which component is failing, how do we measure it, and what evidence shows that the fix worked?”

Microsoft Foundry currently separates RAG evaluation into process measures for retrieval and system measures such as groundedness, relevance, and response completeness. That separation is valuable because a grounded answer generated from weak evidence can still be incomplete, while an apparently complete answer can be ungrounded if the retrieved context does not support its claims.

Amazon Bedrock supports RAG evaluation jobs for retrieval-only and retrieval-and-generation workflows. A team can prepare a query set, test how an Amazon Bedrock Knowledge Base retrieves passages, and compare response quality separately when generation is enabled. Retrieval-only evaluation asks whether the right evidence appeared; retrieve-and-generate evaluation adds whether the model used that evidence well. An evaluator model can assign useful ratings, but human-labeled queries and business-critical examples remain necessary to catch incorrect judgments, outdated policies, and access-control mistakes that a generic quality score might not expose.

In Microsoft Foundry, evaluation can combine built-in quality and RAG evaluators with application-specific criteria and traces. Instead of approving a release solely because average groundedness improved, engineers should segment results by query type, permitted user identity, document freshness, and the presence of conflicting evidence. A deployment can pass easy fact lookups but fail cross-document questions that require a date-qualified exception. Preserve the retrieved document identifiers and resulting model trace for failed cases so the team knows whether to repair indexing, retrieval filters, ranking, prompt construction, or answer-generation behavior.

Evaluate retrieval before judging the answer

The first evaluation layer asks whether the retriever surfaced the documents or chunks that should have been available to the model. If a known answer lives in one policy document but that document never reaches the context window, prompt tuning is unlikely to solve the problem. The failure happened earlier. Retrieval tests should therefore begin with a query set and relevance judgments that identify which sources or passages are expected for each query.

With labeled relevance data, teams can use information-retrieval measures such as recall at k, precision at k, mean reciprocal rank, or normalized discounted cumulative gain. The exact metric matters less than what it reveals. Recall-oriented measures answer whether enough of the needed evidence appears in the candidate set. Precision-oriented measures expose how much irrelevant material is competing for context. Ranking measures show whether the best evidence appears early enough to survive a top-k cutoff.

Without formal relevance labels, qualitative retrieval scoring is still useful, but it should not be confused with end-to-end answer scoring. The engineer should be able to inspect the returned chunks and ask: are they about the right entity, time period, jurisdiction, product version, and level of detail? A high vector-similarity score is not proof that a passage is the right evidence.

Build the relevance labels with enough granularity to support diagnosis. A binary relevant/not-relevant judgment is simple, but graded relevance can distinguish a passage that contains the exact answer from one that is merely useful background. It also helps when multiple documents are valid but one is more authoritative or current. For time-sensitive corpora, the benchmark should state which version of a source is expected so a retriever is not rewarded for returning an obsolete but semantically similar document.

Chunking and embeddings are evaluation variables, not setup details

Chunk design changes the unit that retrieval can return. Very small chunks can isolate exact facts but strip away qualifiers, definitions, and exceptions. Very large chunks preserve context but may dilute the signal used for ranking and consume the context window with material the model does not need. RAG chunking should therefore be evaluated against real queries instead of chosen by a universal token count.

The same principle applies to embeddings. An embedding model that works well for natural-language support questions may behave differently for source code, product identifiers, legal clauses, or multilingual content. Embedding model selection should include representative data and failure cases, especially queries where exact lexical matching competes with semantic similarity.

A useful experiment changes one retrieval variable at a time: chunk strategy, embedding model, metadata filter, candidate count, hybrid weighting, or reranker. Record the configuration with the scores. Otherwise teams can observe a quality change without knowing which change caused it.

Groundedness and completeness answer different questions

Once the right context reaches the model, generation evaluation becomes meaningful. Groundedness asks whether the response is supported by the supplied evidence. Completeness asks whether the response covers the important information it was expected to cover. A concise answer can be perfectly grounded but incomplete; a long answer can look comprehensive while introducing unsupported claims.

Relevance adds another dimension: whether the response actually addresses the user’s question rather than merely paraphrasing the retrieved documents. The three measures often need to be read together. High groundedness with low relevance can indicate that the answer stayed faithful to the context but followed the wrong evidence. High relevance with poor groundedness can indicate that the model answered plausibly from its own knowledge instead of from the approved source set.

For regulated or high-stakes use cases, a team should define what “grounded” means at the claim level. If a response combines a supported fact with an unsupported recommendation in the same sentence, a coarse answer-level score may hide the defect. Citation checks and claim-to-source tracing can make the evaluation more actionable.

Use slices to expose failures hidden by averages

A single aggregate score can conceal important regressions. RAG test sets should be sliced by query type: factual lookup, comparison, multi-document synthesis, policy exception, recency-sensitive question, acronym-heavy query, ambiguous entity, or adversarial wording. A system that averages well may still fail consistently on one class of question that matters to the business.

Data slices matter too. Measure short versus long documents, scanned text versus native text, tables versus prose, frequently updated versus stable sources, and permission-restricted versus public material. A retrieval configuration that performs well on clean documentation may struggle with nested headings, boilerplate, or duplicated versions.

Production incidents should feed the test set. When a real user exposes a failure, preserve the query, expected evidence, observed retrieval trace, and corrected outcome. The evaluation corpus then grows around the system’s actual risk rather than around only synthetic examples.

Measure the cost of getting the right answer

Quality metrics are necessary but incomplete because RAG systems have operational budgets. Larger candidate sets may improve recall while increasing reranking cost and latency. More retrieved chunks may improve completeness while making generation slower and raising token usage. Expensive evaluators can be valuable during release testing but impractical on every production request.

Track end-to-end latency together with stage-level timing. Retrieval, reranking, prompt assembly, model inference, and post-processing should be separately observable. The same applies to token volume and model calls. This lets teams distinguish a slow search index from a slow generation model or an overly aggressive multi-query strategy.

Evaluation should therefore identify a quality frontier rather than a single maximum score. A configuration that improves groundedness by a tiny amount but doubles p95 latency may be the wrong production choice. The acceptable tradeoff depends on the workload and the consequence of error.

Safety and permissions belong in the RAG test plan

Retrieval quality is not only about relevance. A system can retrieve highly relevant content that the user is not authorized to see. Permission filtering should be tested with users, groups, document labels, and revoked access, including cases where the same topic exists in both public and restricted sources. The evaluation should fail if an unauthorized chunk reaches the model even when the final answer does not quote it.

Indirect prompt injection creates another retrieval-specific risk because malicious instructions can live inside retrieved content. Evaluation sets should include documents that attempt to override system instructions, request secret disclosure, or redirect tool use. The system should treat retrieved text as data, not as trusted control instructions.

AI governance turns evaluation criteria into release authority: named owners define acceptance thresholds, approve exceptions, and decide what evidence must be retained. Those controls keep a failed or ambiguous evaluation from becoming an informal engineering judgment with no accountable decision record.

Build release gates from multiple signals

A practical release gate combines retrieval and generation measures with business-specific checks. For example, a release might require a minimum retrieval recall on a labeled benchmark, no material regression in groundedness or completeness, zero permission-boundary failures, acceptable citation precision, and latency within a defined service objective.

Release evaluation should include paired comparisons against the current production configuration. Run the same test cases through old and new retrieval stacks, inspect cases that changed, and classify the reason: candidate recall, rank order, context truncation, model behavior, or policy filtering. That practice makes the decision about shipping a change much stronger than comparing two unrelated aggregate scores.

Thresholds should be compared with a stable baseline. Model-based evaluators are themselves probabilistic, so a small score movement may not be meaningful. Repeated runs, confidence intervals, human review of changed cases, and deterministic retrieval metrics help distinguish noise from a real regression.

The most valuable output of RAG evaluation is not a leaderboard number. It is a diagnosis that tells the team whether to change the corpus, parsing, chunking, embeddings, filtering, retrieval, reranking, prompt, model, or policy layer. That is how evaluation becomes an engineering control instead of a demo score.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!