Hallucination testing is more useful when it is built around concrete failure cases than when it relies on a vague instruction to “check factuality.” Microsoft Foundry supports reusable evaluation datasets, model and agent evaluators, synthetic test generation, trace-derived data, and CI/CD-oriented reruns. For Microsoft AI Agents, that makes a hallucination test set an engineering asset that should evolve alongside the prompt, retrieval layer, tools, and model.
The term hallucination covers several different behaviors: unsupported claims, wrong use of retrieved context, invented tool results, incorrect attributes, and confident answers where the system should abstain. A strong test set labels those failure families so the evaluation stack can use the right metric for each one.
Start with representative production questions, then add adversarial cases
A test set should include the common tasks users actually perform, otherwise a system can score well while failing its main workload. Start with high-volume and high-value query classes, then add cases that historically caused errors or that probe known weaknesses such as ambiguous names, stale documents, conflicting sources, or insufficient evidence.
Generative AI evaluation pipelines become more meaningful when the dataset mirrors real distribution rather than a collection of easy demos. Keep enough difficult cases that an improvement must earn its score instead of benefiting from an overly forgiving test mix.
Encode the expected evidence, not just the expected wording
For RAG tasks, storing only a reference answer can encourage superficial text matching. Preserve the grounding context or retrieval ground truth that should support the response. This allows evaluation to distinguish a semantically different but supported answer from a fluent answer that invented facts.
Enterprise RAG chunking affects what evidence is available, so the dataset should make retrieval assumptions visible. If a test is supposed to retrieve a policy clause, record the expected document or passage rather than treating retrieval as an invisible preprocessing step.
Use multiple evaluators for different hallucination modes
Foundry’s built-in evaluators include groundedness, relevance, response completeness, retrieval metrics, tool-use metrics, and risk evaluators such as Ungrounded Attributes. No single metric captures every factual failure. A response can be grounded but incomplete, relevant but based on the wrong document, or tool-correct but semantically misleading.
LLM judges and evaluation metrics should therefore be assembled into a scorecard that matches the application. For a RAG answer, groundedness and retrieval may dominate. For an agent, tool-input accuracy and tool-output utilization may be equally important because fabricated intermediate state can enter through the action loop.
Make abstention cases first-class tests
Some of the most important hallucination tests are questions the system should not answer. Include missing-data scenarios, unauthorized requests, contradictory evidence, unsupported future predictions, and questions outside the configured knowledge boundary.
This connects to autonomous agent security because the safe outcome is sometimes refusal, clarification, or escalation rather than a best-effort guess. An evaluation suite that rewards answer completeness without testing abstention can accidentally train teams to optimize for confident fabrication.
Version the dataset with the product surface
Foundry evaluation datasets are designed to be reusable across model, prompt, and agent versions. Treat them like test code: assign versions, document changes, keep stable core cases, and add regression examples when incidents occur. Avoid silently replacing difficult cases with cleaner ones because that destroys comparability.
Prompt and model versioning should reference the dataset and evaluator configuration used for approval. A release record is far more useful when it says “passed dataset v12 at threshold X” than when it says “looked good in manual testing.”
Use synthetic generation to expand coverage, not to replace curated truth
Microsoft Foundry can generate synthetic evaluation data and simulated conversations. That is useful for exploring edge cases, varying phrasing, and creating broader scenario coverage. Synthetic cases should still be reviewed before they become a long-term quality gate because generated ground truth can carry its own errors or blind spots.
Curated human-authored cases remain especially important for policy, compliance, specialized technical domains, and business logic where the evaluator model may not know the correct answer independently. Synthetic generation is a coverage multiplier, not a substitute for authoritative expected behavior.
Convert production traces into regression cases carefully
Trace-derived datasets can capture realistic prompts and failure patterns that designers did not anticipate. Before promoting production examples into a permanent test set, remove or transform sensitive data, stabilize external dependencies, and record the context that made the original failure meaningful.
Agent analytics and monitoring can identify clusters worth turning into regression tests: repeated unsupported citations, a tool parameter that is frequently corrected, or a query class with unusual abstention rates. That creates a direct path from observability to prevention.
Use thresholds and slice analysis instead of one average score
Averages can hide severe weaknesses in small but important cohorts. Track performance by scenario family, data source, language, risk class, and workflow type. A release should not pass merely because excellent results on easy questions compensate for failures on a regulated or high-value slice.
LLM regression testing should combine aggregate thresholds with hard blockers for critical cases. Some examples deserve a binary expectation—never invent a customer account number, never claim a tool succeeded when it failed—even if the overall score remains high.
Turn every confirmed hallucination into a better system test
The Microsoft AI-103 production mindset is strongest when evaluation changes the engineering process. A confirmed incident should lead to root-cause analysis: retrieval miss, stale source, prompt ambiguity, model limitation, tool misuse, or application bug. Then add a test that will fail if the defect returns.
Over time, the hallucination test set becomes a record of what the system has learned about its own failure modes. That is more valuable than chasing a universal hallucination percentage. The objective is a release gate that reflects real risk, a dataset that grows from evidence, and an evaluation loop that makes each production failure less likely to recur.
Test-set governance should include provenance for every expected answer. Record whether the ground truth came from policy, a database snapshot, product documentation, a domain expert, or a prior validated interaction. When that source changes, the corresponding test may need to change too. Otherwise a model can appear to regress because the evaluation set itself became stale.
Temporal cases deserve explicit versioning. Questions such as current pricing, model availability, regulatory limits, or product features can change even when the model is behaving correctly. Freeze the intended date or source version in the test metadata, and separate “answer as of this snapshot” from “retrieve the latest answer” scenarios.
For tool-using agents, include tests where the tool returns empty, partial, contradictory, and error results. Hallucinations often appear after a failed dependency when the agent tries to be helpful by filling the gap. The expected behavior may be to retry, choose another source, ask for clarification, or state that the answer cannot be verified.
Use canary cases in production after release. A small set of safe, automated prompts can verify that critical grounding and abstention behaviors still work after infrastructure changes, model updates, or connector changes. Canary results do not replace offline evaluation, but they can detect environmental drift that a static pre-release test could not anticipate.
Finally, periodically prune redundant tests while preserving failure diversity. A suite with thousands of near-identical cases becomes expensive without adding much signal. Keep representative examples for each root-cause family, retain high-severity incidents permanently, and use sampling to cover natural language variation. The quality of the test taxonomy matters more than raw dataset size.
Balance the dataset between exact-fact tests and reasoning tests. Exact facts reveal unsupported invention around names, numbers, and dates; reasoning tests reveal whether the model draws conclusions the evidence does not justify. Both matter because a system can copy facts accurately yet hallucinate the relationship between them.
Store expected failure handling as part of the test case. Some cases should expect a direct answer, others a clarification question, a refusal, a tool call, or an explicit statement that evidence is insufficient. This makes the evaluation suite reflect the product contract rather than assuming every test should end with a factual paragraph.
Measure evaluator stability too. Model-based judges can change when their underlying deployment or prompt changes, so periodically rescore a frozen calibration set and compare results. If the judge drifts, release thresholds may need recalibration. Quality gates are only trustworthy when the measurement instrument itself is monitored.
Keep a small manually reviewed calibration subset that domain experts revisit on a schedule. This anchors the automated metrics to human expectations and helps detect when ground truth, evaluator behavior, or product requirements have changed. The calibration set should contain both obvious cases and difficult borderline examples.
Test-set maintenance needs the same discipline as application data. Record the source of each case, expected behavior, acceptable evidence, severity, domain, and the reason the example was added. When a production incident becomes a regression test, capture the failure mode rather than copying only the exact wording; otherwise a model may pass a near-duplicate without demonstrating that the underlying weakness is fixed. Maintain slices for unsupported factual claims, fabricated citations, incorrect synthesis across sources, refusal failures, stale knowledge, and overconfident uncertainty. Release decisions can then examine both the aggregate score and the high-severity slices. A small improvement in average performance should not hide a regression in the specific hallucination class that carries the greatest business or safety consequence.