AI red teaming is an adversarial evaluation practice designed to uncover harmful behaviors, security weaknesses, misuse pathways, failure modes, and gaps that ordinary testing may miss. Governance determines which systems are red-teamed, who performs the testing, how realistic attacks can be, what data or users may be involved, how findings are handled, and which results can block deployment.
Within AI Governance, red teaming is one source of risk evidence. NIST’s 2026 ARIA Evaluation Planning Manual describes holistic AI evaluation through a combination of model testing, red teaming, and user testing, reinforcing that adversarial testing belongs inside a broader evaluation program.
Red teaming creates value only when findings influence engineering and release decisions.
Scope should begin with the system, not only the model
A model-only test can miss retrieval, tools, permissions, memory, human workflow, and application logic that create real risk.
Define whether the exercise targets the base model, prompt, RAG pipeline, agent, tools, user interface, policy layer, or full production-like system.
The scope should match the harm scenarios the organization actually cares about.
Rules of engagement protect people and systems
Red teams may test unsafe content, security bypass, sensitive data exposure, or destructive tool behavior.
Define allowed targets, prohibited actions, test accounts, data boundaries, escalation contacts, time windows, and stop conditions.
The exercise should be adversarial toward the system without becoming uncontrolled toward real users or production infrastructure.
Independence should be proportional to risk
Developers can red-team their own system and often find useful issues, but they carry assumptions about how the design is supposed to work.
Higher-risk systems may justify an independent internal team or external assessors who do not share the same design assumptions.
Independence is not binary; governance should define how much separation is appropriate for the system’s impact.
Threat models should guide test scenarios
AI red teaming can cover prompt injection, data leakage, harmful content, tool abuse, privilege escalation, policy bypass, hallucination, impersonation, fraud enablement, or unsafe automation.
Prioritize scenarios based on system capabilities and affected people rather than one generic “jailbreak list.”
A read-only summarizer and a payment agent have very different attack surfaces.
Test evidence should be reproducible
Record model/application version, prompt, data state, tool permissions, red-team input, observed output, environment, and expected behavior.
Many AI failures are nondeterministic, so one finding may require several repetitions to understand frequency and severity.
Evidence should be sufficient for engineers to reproduce the issue without exposing dangerous details more broadly than necessary.
Severity should consider capability and exploitability
A disturbing response is not necessarily the same severity as an exploitable path to unauthorized payment or sensitive data.
Score findings using impact, likelihood, user exposure, ease of exploitation, required privileges, detectability, and existing mitigations.
Use the organization’s existing AI risk taxonomy so red-team findings integrate with normal risk reporting.
Findings should have remediation owners and deadlines
Every material finding needs an owner, treatment, due date, retest plan, and release impact.
Some findings require model or prompt changes; others require tool permissions, UI changes, monitoring, business process redesign, or narrower product scope.
Red teaming should not become a report repository disconnected from delivery teams.
Release criteria should be defined before the test
If the organization decides after seeing results which findings matter, schedule pressure can distort the decision.
Define severity thresholds, mandatory fixes, allowable risk acceptance, and who can approve deployment with open findings before the exercise begins.
AI Risk Acceptance Decisions covers the residual-risk path.
Retest should validate the mitigation, not only the original prompt
A narrow fix may block one known attack while leaving the underlying weakness intact.
Retest the original case plus variants and neighboring attack strategies.
The goal is evidence that the control reduced the class of risk, not merely that one test string no longer works.
Red teaming should recur after material change
New tools, model upgrades, broader permissions, new data sources, new languages, or a different user population can introduce new attack paths.
Change control should therefore trigger targeted red-team regression where the threat surface changed.
A two-year-old red-team report is weak evidence for a substantially different production agent.
Governance should protect sensitive findings
Red-team reports can contain exploit chains, unsafe outputs, sensitive data, and control weaknesses. Distribution should be limited to people who need the details.
Portfolio reporting can summarize severity and remediation without exposing the exploit recipe.
Good red-team governance increases transparency to decision-makers while maintaining responsible handling of dangerous findings.
Red-team datasets and attack libraries should be controlled. Some prompts may contain harmful instructions, exploit techniques, or personal data that should not be shared broadly. Store them in restricted repositories with clear retention and access rules.
Tester safety matters as well. Repeated exposure to disturbing generated content can create psychological harm. High-intensity exercises should provide opt-out mechanisms, content warnings, rotation, and support appropriate to the material being tested.
Production red teaming needs stronger safeguards than isolated test environments. If real systems must be probed, use dedicated accounts, rate limits, non-customer data where possible, and rollback plans. Deliberately triggering unsafe behavior in production without containment can become the incident itself.
Red teams should include multidisciplinary perspectives when the risk demands it. Security specialists are strong at adversarial access paths; domain experts can identify harmful but technically “correct” outputs; UX researchers can expose user-manipulation risks; privacy teams can identify data misuse.
Findings should be normalized against ordinary vulnerability and risk processes where possible. A red-team issue that reaches an external tool with excessive privilege may belong in the security vulnerability system as well as the AI risk register.
Metrics should avoid gamification. Counting number of prompts tested or number of jailbreaks found does not prove the system is safe. More useful measures include high-severity finding closure, regression coverage, repeated failure classes, and whether mitigations reduce exploitability.
Governance should preserve adversarial independence while keeping the process collaborative enough that findings can be fixed. The goal is not to “beat” the development team; it is to uncover risk before users or attackers do.
Coverage should include ordinary user behavior as well as malicious behavior. Some serious failures occur when normal users phrase ambiguous requests, upload unexpected formats, or rely too heavily on confident output.
Red-team findings should be compared with production incidents and support tickets. If users repeatedly discover a failure class the red team never tested, the threat model should be updated.
Vendor-hosted models may limit what internals can be inspected, but system-level red teaming can still test application controls, retrieval, tool permissions, response handling, and model behavior through the exposed interface.
Testing cadence can be risk-tiered. High-impact, frequently changing agents may need red-team regression after every major capability change, while stable low-impact tools can be tested less often.
The governance objective is sustained adversarial learning: each exercise expands the organization’s test corpus, threat model, mitigation patterns, and understanding of where its AI controls fail under pressure.
Red-team environments should match production capabilities closely enough that findings are meaningful. Testing a model without the tools, permissions, memory, retrieval, and guardrails used in production can miss the interactions that create the highest risk.
At the same time, isolation should prevent test actions from causing real harm. Synthetic identities, sandbox tools, mock payment systems, and scrubbed datasets can preserve realistic behavior without touching live customer resources.
Red-team evidence should be incorporated into regression suites after remediation so the organization does not repeatedly rediscover the same weakness in future model or prompt versions.
Finding disclosure should be coordinated with remediation. Broadly publishing an exploitable agent weakness before controls are fixed can increase risk, while hiding all findings indefinitely prevents organizational learning. Use audience-appropriate summaries and restricted technical detail.
Red-team exercises should also test recovery. If an unsafe action is triggered, can the system stop, roll back, revoke credentials, remove malicious memory, and alert an operator? Resilience after exploitation is part of system safety.
Results should influence future test planning. Repeated weaknesses in retrieval, tool authorization, or prompt boundaries should receive more coverage in the next exercise rather than restarting from a generic checklist.
Metrics should include time to remediate high-severity findings and percentage retested successfully. A red-team program that finds the same serious issues repeatedly but cannot close them is not improving resilience.
Governance should also define when external red teaming is required—for example, especially novel, high-impact, or regulator/customer-sensitive systems where independent evidence has additional value.
The mature program makes adversarial evaluation repeatable, safe, independent enough to challenge assumptions, and connected directly to release authority and long-term regression testing.
Exercise planning should include success criteria for the red team itself: which capabilities, user paths, tools, languages, and risk scenarios must be covered before the evidence is considered sufficient for the release decision.
That coverage record should survive the exercise so future teams know which attack classes were actually tested and which remain assumptions. Red teaming is credible when decision-makers can see both the findings and the boundaries of what the exercise did not examine.
Preserve that scope statement with the release evidence so future reviewers understand exactly what confidence the test can support.