Red-teaming an AI application is not the same exercise as asking a model a collection of provocative prompts. A production system can retrieve documents, call tools, write records, trigger workflows, and expose different data to different users. The useful test case therefore has to describe an adversarial situation in which a model decision interacts with application authority, data, and downstream controls.
For applications built on AWS, AIP-C01 relates red-team testing to the exam’s AI safety, governance, and testing objectives: a failure is significant when it crosses an actual boundary around tool execution, retrieval permissions, or sensitive data. On Azure, AI-103 measures related skills in Microsoft Foundry, including risk evaluation, safeguards, trace-based oversight, and agent tool controls. These are overlapping responsibilities, not evidence that the two certification exams contain identical red-team objectives. A credible red-team case therefore identifies the user and service identities, accessible tools, source of hostile input, prohibited outcome, and observable pass/fail criteria before anyone writes an attack prompt.
Start with a threat hypothesis, not a list of clever prompts
A strong test begins with an asset and a failure hypothesis. The asset might be confidential retrieval content, an administrative tool, a payment action, a system prompt, a customer record, or an approval workflow. The hypothesis then states how an attacker could cross the intended boundary: malicious user input might alter tool arguments, retrieved content might inject instructions, a model might expose data from another tenant, or an output might become executable markup downstream.
This makes the case reproducible. Instead of “try to jailbreak the assistant,” the test can state that an untrusted document contains instructions asking the model to ignore the user’s request and send retrieved secrets to an external destination. The preconditions, attacker-controlled input, expected control, observable evidence, and pass/fail condition can all be recorded. That structure is also what turns AI red-team governance into an engineering discipline: test evidence can be assigned, reviewed, compared across releases, and linked to the risk that justified the test.
Separate model behavior from system impact
A model can produce an undesirable sentence without creating material harm, while a seemingly harmless response can trigger a high-impact tool call. Red-team cases should therefore distinguish model-level behavior from system-level consequence. At the model layer, teams may examine instruction following, refusal boundaries, hallucination, sensitive-data reproduction, and manipulation. At the application layer, they should examine authorization, data access, tool invocation, persistence, network egress, rendering, and workflow side effects.
For agentic systems, the test must separate model intent from downstream effect. A model saying “I will delete the account” is different from an application actually issuing a delete request with valid credentials. The agentic AI engineering control model treats autonomy as a sequence of governed decisions, so a red-team case should record the proposed action, authorization decision, tool invocation, and resulting state separately. That evidence identifies which boundary failed instead of treating the model as the only security perimeter.
Indirect prompt injection needs hostile retrieval fixtures
Direct attacks are easy to stage because the attacker controls the conversation. Indirect attacks are more representative of retrieval and browsing systems because the instruction is embedded in content the application considers data. A poisoned support ticket, wiki page, PDF, issue comment, web result, or tool response can contain language intended for the model rather than the human reader.
A practical fixture should preserve that distinction. The user request can be benign, the retrieved document can contain the attack, and the expected behavior can require the application to treat the retrieved instruction as untrusted content. Tests for indirect prompt injection should vary where the malicious text appears, how strongly it conflicts with trusted instructions, what tools are available, and whether sensitive context is present. The objective is not to prove that one attack string is blocked. It is to learn whether the control survives changes in wording, document position, retrieval rank, and task context.
Tool tests should target authorization and argument integrity
Once an agent can call tools, red-team design should include actions that are valid in shape but invalid in authority. Examples include asking a read-only user to invoke an administrative operation, modifying an identifier after approval, expanding the scope of a query, substituting an attacker-controlled destination, or using one tool’s output to construct unsafe arguments for another.
These cases expose a common architectural error: allowing model intent to substitute for policy. The model may select a tool and propose arguments, but the application still has to authenticate the caller, authorize the operation, validate parameters, constrain destinations, and enforce transaction rules. A passing test therefore should not depend solely on the model refusing. It should show that deterministic controls stop an unauthorized operation even if the model attempts it.
Tool descriptions themselves are part of the attack surface. Ambiguous schemas, overly broad operations, and tools that combine read and write behavior make failures harder to contain because a compromised planner can express more authority through a single call. Clear MCP tool contracts narrow operations, constrain inputs, and give the red-team harness specific authorization and validation points to probe.
Data-exposure tests need identities, scopes, and canary records
“Does the model leak data?” is too broad to be a useful case. A better test specifies which identity is active, which corpus or tenant the identity may access, what sensitive field is planted as a canary, and which path could expose it. The test may attempt cross-tenant retrieval, prompt-based requests for hidden instructions, inference from prior conversation context, logging exposure, or exfiltration through a tool.
Canary data makes the result observable without relying on subjective review. Synthetic secrets, unique document markers, or deliberately inaccessible records can show whether a boundary was crossed. The same technique helps test retrieval filters: the system can contain documents that are semantically attractive to the query but forbidden to the active principal. A secure result should preserve authorization even when the forbidden document is the closest vector match.
Safety tests need both attack success and legitimate-task preservation
A defense that blocks every difficult request can look impressive in an attack-only dataset while making the application unusable. Red-team suites should therefore contain paired legitimate cases. If a control rejects a malicious request to export credentials, the suite should also contain a permitted administrative export with the right identity and approval. If a content filter blocks unsafe instructions, the suite should include legitimate discussion of the same vocabulary in educational, compliance, or incident-response contexts.
A mature red-team suite also measures regression behavior rather than recording isolated attacks. Regression testing can track both unsafe passes and legitimate-task rejections across model, prompt, retrieval, and tool changes. Measuring both directions prevents a control from appearing successful merely because it blocks everything.
Expected results should describe observable controls
“The model should be safe” is not an acceptance criterion. A red-team case should name the observable outcome that proves the boundary held. The application might refuse a tool invocation, strip an untrusted instruction from executable context, require human approval, return a constrained schema, omit unauthorized records, sanitize output before rendering, or emit a security event containing the relevant trace identifiers.
An Azure deployment can put attack fixtures into Microsoft Foundry evaluation datasets and compare agent task-completion signals with prohibited actions, sensitive-data leakage, and tool-call accuracy. For example, if an adversarial search result attempts to convince a support agent to retrieve another tenant’s records, the case should fail on unauthorized retrieval or an attempted write—not merely on the presence of an unsafe phrase in the answer. The trace needs to identify the source document, principal, selected tool, arguments, policy decision, and resulting side effects; identifiers should be protected so test logs do not become another data leak.
An AWS deployment can evaluate the same risk at a different enforcement layer: configure Amazon Bedrock Guardrails for supported prompt-attack and sensitive-information checks, then exercise a Bedrock-powered application’s IAM and application authorization checks with a restricted test identity. A blocked guardrail finding is useful evidence, but it does not prove that a tool result, backend API, or every retrieval chunk received the same screening. The test must separately demonstrate that a disallowed operation could not execute even when the model proposed it, and it should measure whether legitimate user workflows still succeed.
The expected result should also say what must not happen. No external request should be sent, no privileged tool should execute, no secret should appear in the response or logs, and no durable state should change before approval. These negative assertions matter because an application can return a safe-looking chat message after an unsafe side effect has already occurred.
Severity should follow consequence, not prompt novelty
A spectacular jailbreak that produces embarrassing text can be less important than a mundane argument-manipulation path that changes a bank account or deletes a production resource. Red-team triage should use consequence, exploitability, required access, affected population, detectability, and recoverability. That keeps engineering attention on material failure modes rather than the most theatrical transcript.
The governance layer should connect findings to owners, risk statements, remediation dates, retests, and release decisions. AI governance supplies that accountability chain so a failed red-team case becomes an owned production risk with evidence of closure rather than an isolated security observation.
Multi-turn and multi-agent tests should preserve attacker state
Many failures do not appear in a single turn. An attacker may first persuade the system to store a misleading preference, then invoke a tool later when the original manipulation is no longer visible. Memory features can preserve hostile content across sessions; delegated agents can pass untrusted instructions between components; a planner can create a benign-looking subtask whose output becomes dangerous only when another agent consumes it. Red-team fixtures should model these sequences instead of resetting the application after every prompt.
The test record should identify which state survives between steps and which trust boundary each component assumes. If one agent retrieves content and another has write authority, a case should check whether the second agent can distinguish user intent from instructions embedded in the first agent’s result. Sequence-aware testing is especially important because each individual message can look acceptable while the combined path crosses a policy boundary.
Red-team suites should become release assets
Adversarial testing is most valuable when important failures become durable regression cases. After a vulnerability is fixed, the triggering fixture and a set of nearby variations should remain in the suite. Model changes, prompt changes, retrieval tuning, new tools, and orchestration changes can all reintroduce a behavior that previously appeared resolved.
Teams should version test inputs, environment assumptions, model configuration, tool availability, expected outcomes, and scoring logic. They should also retain enough trace data to explain a failure rather than recording only a final pass/fail label. The result is a security test library that grows with the system’s authority. Red-teaming then stops being a one-time exercise before launch and becomes a repeatable way to challenge the exact boundaries that make an AI application safe to operate.