Microsoft AI-103: Safety Evaluation in Microsoft Foundry

Safety evaluation is the discipline of testing whether an AI application behaves acceptably before the same failure becomes a production incident. In Microsoft Foundry, this goes beyond checking a chatbot for rude language. The current evaluation stack covers content risks, agent behavior, tool use, groundedness, task completion, and other dimensions that can be run against models, applications, or agents. For teams working through Microsoft AI Agents, the practical goal is to turn safety from a subjective launch review into a repeatable release gate backed by representative datasets and clearly defined thresholds.

Microsoft’s Foundry evaluation service includes hosted risk and safety evaluators for violence, sexual content, self-harm, hate and unfairness, protected material, indirect prompt injection, code vulnerability, and ungrounded attributes. Agent-focused evaluation extends the picture with measures for prohibited actions and sensitive data leakage, alongside task adherence, task completion, intent resolution, tool selection, tool-input accuracy, and tool-call success. That breadth matters because an agent can produce harmless prose while still taking the wrong action. Safety evaluation therefore has to inspect both what the system says and what it does.

Start with a risk model instead of choosing evaluators by convenience

The strongest evaluation plan begins with the application’s actual failure modes. A public information assistant may care heavily about harmful content, misinformation, and accessibility. An internal operations agent may care more about authorization, sensitive-data leakage, prohibited actions, and incorrect tool parameters. A coding assistant adds code vulnerability and license-related concerns. Selecting every available evaluator without mapping it to a real risk creates dashboards that are busy but difficult to act on.

AI guardrails and content safety help define where prevention controls operate, while evaluation tells you whether those controls and the rest of the application achieve the intended result. The distinction is important: a runtime filter can block one category of unsafe output, but an evaluation suite can discover whether the workflow still leaks data through tools, follows malicious context, or fails a user task after the filter intervenes.

Build test datasets around realistic conversations and failure boundaries

A useful safety dataset contains ordinary requests, edge cases, adversarial prompts, ambiguous requests, and benign examples that resemble attacks. If the dataset contains only obvious jailbreak phrases, it measures the system against an artificial threat model. Include multi-turn conversations where intent changes, retrieval results that contain misleading instructions, tool calls with missing or malformed parameters, and cases where the correct behavior is to ask for clarification rather than act.

Generative AI evaluation pipelines are strongest when each case records the user intent, expected behavior, relevant policy, acceptable variations, and the evaluator that will judge the result. This makes failures diagnosable. A low score should lead to a concrete question such as whether the prompt, retrieval source, tool schema, model, or policy needs adjustment instead of becoming an unexplained red cell on a dashboard.

Evaluate content risk and agent behavior as separate dimensions

Content-safety scores and agent-behavior scores answer different questions. A response can avoid hateful or violent content yet still violate policy by sending an unauthorized message, querying the wrong customer record, or exposing a secret in a tool argument. Conversely, an agent can follow every workflow rule while generating content that is inappropriate for the audience. Treating one dimension as a proxy for the other creates false confidence.

Autonomous agent security should therefore be tested with scenarios that include the tools and permissions the production agent actually receives. Foundry’s agent evaluators can inspect process-level behavior such as tool selection and parameter accuracy, while risk evaluators inspect output and agent-specific risks. The release decision should combine these signals according to business impact rather than collapse them into one average score.

Use indirect-attack evaluation for retrieval and tool-fed context

Indirect prompt injection is especially important for agents that read documents, email, tickets, websites, or tool output. The user may ask a safe question while the retrieved content contains instructions that attempt to redirect the model. Foundry’s indirect-attack evaluator is designed to measure whether model behavior was altered by this injected context. Test cases should preserve the source boundary so the system can distinguish user intent, system policy, and untrusted evidence.

Prompt Shields for AI apps can reduce runtime exposure, but evaluation still needs cases where the attack is missed, partially detected, or appears in an unexpected channel. A resilient design assumes at least some hostile content will reach the model and verifies that authorization, tool restrictions, citations, and approval controls prevent a missed detector signal from becoming a harmful action.

Test tools by intent, selection, arguments, and outcome

Agent evaluation should not stop after confirming that a function call was syntactically valid. The selected tool must be appropriate for the user request, arguments must be correct and authorized, the tool result must be interpreted accurately, and the final response must reflect the real outcome. A well-formed call to the wrong system is still a failure. So is a correct call whose error response is ignored.

Agent tools and multi-step reasoning provide a useful mental model for process evaluation. Build cases where two tools appear plausible, where a required field is missing, where a tool returns incomplete data, and where the correct next action is to stop. These cases reveal whether the agent is reasoning within its operating contract rather than simply demonstrating that it can invoke functions.

Pin the variables that must remain comparable across evaluation runs

Evaluation results are hard to interpret if the dataset, model deployment, evaluator version, prompt, rubric, and tool definitions all change simultaneously. Microsoft recommends pinning relevant datasets, deployments, evaluator versions, and rubric versions when results must be comparable. A release gate should make those dependencies explicit so teams can tell whether a score changed because the application improved or because the measuring instrument changed.

Agent lifecycle management should treat evaluation configuration as a versioned production asset. Store the test-set version and evaluation policy with the release candidate, record exceptions, and retain enough evidence to reproduce the decision. This is especially important when preview evaluators evolve or when model upgrades change behavior without an application-code change.

Turn evaluation thresholds into release gates with clear exception rules

Thresholds should reflect risk, not aesthetic preference. A content assistant may tolerate a small decline in stylistic fluency but not an increase in sensitive-data leakage. A high-impact workflow may require zero successful prohibited-action cases in a critical test suite even if overall task-completion score remains excellent. Define pass criteria by category and severity, then state who can approve an exception and how long that exception remains valid.

Governance standards and procedures make this process auditable. The release record should show which tests ran, which failed, what mitigation was accepted, and who owned the residual risk. This keeps evaluation from becoming a ceremonial report that everyone reads after deployment rather than a control that can genuinely stop a risky release.

Use traces and production incidents to improve the test set

No predeployment dataset captures every real conversation. Production traces, support tickets, user corrections, blocked actions, and near misses should feed a controlled process for adding new regression cases. The important step is to sanitize and classify these examples before reuse so evaluation datasets do not become a repository of unnecessary personal or confidential data.

Agent analytics and monitoring connect operational evidence back to evaluation. When a production issue appears, reproduce the relevant trajectory in a test environment, add a case that would have caught it, fix the system, and verify the case stays passing. This creates a feedback loop in which the safety suite becomes more representative over time rather than remaining frozen at launch.

Measure safety continuously because the system keeps changing

Model updates, new tools, revised retrieval indexes, changed prompts, broader permissions, and new user populations can all shift risk. Re-run the relevant evaluation suite after meaningful changes and on a periodic schedule for long-lived agents. For high-risk workflows, combine deterministic checks, model-based evaluators, human review, and targeted red-team exercises rather than relying on one scoring method.

Microsoft Foundry provides a growing set of evaluators and managed reporting, but the quality of the result still depends on the scenarios and decision rules supplied by the team. Safety evaluation works when it reflects the real operating boundary: the users, data, tools, permissions, and consequences the agent will encounter. The target is not a perfect score. It is evidence that known high-impact failure modes are understood, tested, and controlled before production traffic exposes them.

Teams should also distinguish a safety test from a compliance determination. An evaluator can provide evidence that a system behaved a certain way on a defined dataset; it does not prove that every regulatory or organizational requirement has been met. Keep legal, privacy, accessibility, security, and domain-specific review connected to the evaluation program, but do not collapse them into one automated score. This distinction makes the evidence more credible because each control has a clear purpose and owner.

For mature programs, maintain a small critical suite that runs on every release and a broader suite that runs on scheduled or high-risk changes. The critical suite should cover the failures that would stop deployment immediately, while the broader suite explores diversity, long-tail behavior, and emerging risks. This tiered structure keeps feedback fast enough for engineering without sacrificing deeper periodic assurance.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!