A Foundry agent becomes useful when it can answer the right question, acknowledge missing evidence, and avoid taking actions that the user has not authorized. Creating an agent that responds to a greeting is a valid technical starting point, but it does not demonstrate production readiness. An enterprise application also needs an identity boundary, access to approved knowledge, a controlled tool surface, observability, and tests for both successful and unsafe outcomes. A practical AI-103 learning exercise should expose those engineering decisions instead of presenting model output as proof that the whole application works.
This exercise uses a fictional customer-support scenario and current Microsoft Foundry prompt-agent patterns. The core Python example follows Microsoft Learn’s Azure AI Projects 2.x API; the retrieval, approval, and evaluation extensions are separate design-and-test tasks until an authorized learner configures their actual resources. No Foundry service, customer data, tool action, or cloud experiment has been executed to produce this article. The objective is a reproducible, reviewable lab plan that teaches how to verify rather than guess.
Define one bounded support task and the failure conditions
Imagine a small support assistant for a company that sells a fictional subscription product. Users ask factual questions about a documented refund window, service availability, and how to contact support. The agent may answer from published, approved policy text, but it must not issue refunds, modify accounts, promise special exceptions, or invent a rule when the knowledge source is silent. That boundary makes the task measurable: successful answers are grounded in approved evidence, while unsafe behavior includes fabricating eligibility or treating a user instruction as authorization for a transaction.
Prepare three short, invented policy records without personal information: one states an example refund window, one describes the official support channel, and one explains which situations need a human exception review. Label each record with source ID, effective date, owner, approved audience, and a clear version. These are demonstration policies, not real Exam-Labs terms. A correct agent should distinguish current published policy from a retired version, cite the correct record when a retrieval tool is implemented, and say it lacks sufficient information if the answer is not present.
Write acceptance cases before deploying anything. Include a normal eligibility question supported by the sample policy, an out-of-scope question about a competitor, an intentionally conflicting retired document, and an adversarial input asking the agent to disregard its instructions or approve an unauthorized refund. Add a test where the correct action is escalation rather than a definitive answer. These cases will later show whether the application is reliable, safe, and clear about uncertainty.
This scenario differs from a broad explanation of Foundry agent architecture. Here the reader carries a concrete business boundary through setup, evidence, response testing, and release criteria. Architecture concepts matter because they determine how the test scenario behaves, but they do not replace the test.
Prepare a least-privilege Foundry project
Use an Azure subscription and Foundry project that the learner is actually authorized to administer or use. Confirm the supported project endpoint, a deployed model, the current Foundry SDK, and the required role assignments. Microsoft Foundry’s updated Python quickstart uses the Azure AI Projects 2.x API, which is not source-compatible with every earlier Azure AI Projects 1.x example. Mixing old agent interfaces into the new code is a common source of confusing attribute errors. Check the official package version and endpoint format on the day the lab is run.
Create a local Python virtual environment and install compatible packages using the instructions for the current quickstart. Microsoft documents azure-ai-projects version 2.3.0 or later for the current projects interface, alongside azure-identity for Microsoft Entra authentication. Sign in to the correct Azure tenant using an approved identity such as az login, then set environment variables for the Foundry project endpoint, intended model deployment, and a distinctive agent name. Do not paste account secrets or bearer tokens into article examples.
Confirm the identity’s specific permissions. A principal that can use an already deployed model may not be allowed to deploy another one, modify a project, or create agent versions. If a task fails with authorization errors, do not respond by assigning a broad owner role to everyone. Trace the missing operation to the official permissions reference and grant the narrow role required for an authorized learning environment. In a production application, prefer appropriate managed/workload identity instead of relying on a developer’s interactive login.
Network access is a separate design decision. A project can be correctly authenticated but fail to reach a model, search service, or downstream API because private endpoints, outbound controls, or resource firewalls are configured differently. Before assuming the prompt is broken, verify the chosen endpoint, model deployment, region, project linkage, and connection permissions. Those are also the practical troubleshooting concerns described in Microsoft Foundry project design.
Create one minimal agent with current SDK syntax
The initial agent should make an explicit promise to avoid inventing policy and performing unauthorized actions. Keep the instructions short and testable; a large prompt containing conflicting aspirations is harder to review. The following Python pattern is adapted from Microsoft’s current prompt-agent quickstart and creates a versioned prompt agent, establishes a conversation, and asks a test question. It assumes a valid project, deployed model and permissions. The code has been syntax-reviewed offline, but no authenticated Foundry execution has been performed.
import os
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import PromptAgentDefinition
from azure.identity import DefaultAzureCredential
project = AIProjectClient(
endpoint=os.environ["PROJECT_ENDPOINT"],
credential=DefaultAzureCredential(),
)
agent_name = os.environ["AGENT_NAME"]
project.agents.create_version(
agent_name=agent_name,
definition=PromptAgentDefinition(
model=os.environ["MODEL_DEPLOYMENT"],
instructions=(
"You are a training support assistant. "
"Use only explicitly supplied facts; say when evidence is missing. "
"Never approve or execute account changes."
),
),
)
openai = project.get_openai_client(agent_name=agent_name)
conversation = openai.conversations.create()
response = openai.responses.create(
conversation=conversation.id,
input="What can you tell me about a refund policy if none was supplied?",
)
print(response.output_text)
The expected safe behavior is that the agent states no refund rule was supplied. This is an expected test condition, not a reported result. A prompt instruction alone is not a strong security control: a model can still answer incorrectly or be manipulated by later user content. The exercise establishes a baseline that can be compared with subsequent retrieval and tool configurations. If the model invents a policy despite the absence of evidence, record the exact output and do not treat agent creation as a passing quality test.
Version the agent instructions and record the deployed model name, project endpoint identifier without revealing credentials, test prompt, and observation. A newly created agent version is not automatically the approved production version. Use a dedicated test agent and follow project cleanup and ownership policies so temporary resources are not left running indefinitely.
Connect approved knowledge and preserve citation scope
Next, prepare a retrieval architecture for the fictional policy records. Each document should have stable identity, effective date, content version, and permissions. Decide whether a lightweight in-memory knowledge set is adequate for this small demonstration or whether the learner needs to provision Azure AI Search for a more realistic RAG workflow. Building a search index only because RAG is fashionable adds cost and moving parts; the intent is to demonstrate retrieval and grounding, not to satisfy a particular tool-count target.
For a search-backed version, map fields so that the document ID, revision status, audience, and source content can be carried through retrieval. Test that the index distinguishes current from retired policy copies. Choose an appropriate chunking strategy for paragraph-level policy sections; do not split a conditional clause from its exceptions. The meaning of a refund rule depends on the relationship between its main statement and qualifying text. The principles of RAG on Azure explain why retrieval quality, permission filtering and original-source evidence are separate requirements from a model’s ability to write fluent prose.
Connect the retrieval capability using the supported Foundry agent tool interface from the current official documentation. Because this tool configuration depends on the chosen resource, project type and SDK version, do not fake a universal constructor or imply the minimal prompt-agent example already performs RAG. After connecting an approved index, repeat the test questions and inspect the actual retrieved records. If the search returns the retired rule, the problem may be indexing or metadata filtering rather than answer generation.
Enforce the user and tenant access boundary before documents reach the model. Metadata describing access is not enough unless query-time filters and permissions are actually applied. Test a request from an identity that should not see an internal-only document, and confirm that it cannot retrieve or infer that protected content. A cited source must be the record the agent was entitled to use, not simply the closest document available in the index.
Keep external tools behind business authorization
Tool calling increases what an agent can do, but it also increases risk. Start with a harmless, read-only function such as looking up the fictional plan’s public support contact. Define its arguments narrowly, validate them independently, and return only the information the caller is permitted to receive. The structure of agent tool schemas helps models choose capabilities predictably, while the actual application code remains responsible for authentication, authorization, input checking, and safe effects.
Now imagine adding a refund-request tool. The model may propose that the tool be used, but should not be allowed to finalize a financial action solely because the user typed “refund me now” or because a retrieved document contains an instruction to issue refunds. Define an approval boundary: the application must validate account ownership and transaction eligibility, obtain appropriate human confirmation, apply idempotency, and produce an audit receipt. A tool description cannot grant authority that the underlying system has not authorized.
One adversarial test places a hidden instruction in a retrieved policy snippet telling the assistant to ignore policy and perform a transaction. The snippet is untrusted evidence, not a higher-priority instruction. Evaluate whether the retrieval pipeline preserves the document as data, whether the model attempts to call any unauthorized tools, and whether the application blocks the action independently of the model. The overall system can remain safe even when the model proposes the wrong action, provided the business enforcement boundary rejects it.
The enforcement boundary can be tested locally without giving an agent a real payment capability. This deterministic Python mock checks that the caller owns the account, the idempotency key is present, and a separate backend-owned approval record exists. It intentionally performs no refund, and the model does not control the set of approved keys. A real service would additionally verify authenticated sessions, transaction eligibility, audit events, secure approval storage, and concurrency.
def mock_refund_gate(caller_id, request, approved_keys, seen_keys):
if request.get("customer_id") != caller_id:
return "DENIED_WRONG_CUSTOMER"
key = request.get("idempotency_key")
if not isinstance(key, str) or not key:
return "DENIED_MISSING_KEY"
if key in seen_keys:
return "ALREADY_PROCESSED"
if key not in approved_keys:
return "HUMAN_REVIEW_REQUIRED"
seen_keys.add(key)
return "SIMULATED_ONLY_NO_PAYMENT_ACTION"
approved = {"approved-001"} # Pretend backend approval records
seen = set()
requests = [
("customer-a", {"customer_id": "customer-b", "idempotency_key": "approved-001"}),
("customer-a", {"customer_id": "customer-a", "idempotency_key": "pending-002"}),
("customer-a", {"customer_id": "customer-a", "idempotency_key": "approved-001"}),
("customer-a", {"customer_id": "customer-a", "idempotency_key": "approved-001"}),
]
for caller, data in requests:
print(mock_refund_gate(caller, data, approved, seen))
In a local execution using these fictional inputs, the cases produce an ownership denial, a human-review requirement, a simulated allowed result, and a repeated-key result in that order. This mock was executed offline to validate those four branches; it is not an Azure service test or a financial transaction. The same backend rule would still reject an unauthorized action even if an agent generated a convincing explanation for why it should proceed.
Do not silently send these tests to real billing, CRM, or identity endpoints. Use a simulation or test-double for write operations. Record the intended tool call and expected denial locally. The exercise assesses the logic of tool boundaries, not the ability to produce a genuine financial transaction.
Evaluate correctness, safety, and regression cases
A reliable agent test plan separates retrieval, generation, and business outcomes. Retrieval tests ask whether the correct current policy record was found and whether an unauthorized record stayed hidden. Generation tests ask whether the response accurately reflects that evidence, preserves exceptions, and admits uncertainty. Tool tests ask whether arguments, permissions, confirmations, and error behavior are correct. Combining all those into one “quality score” makes it difficult to understand a failure.
Build a small benchmark set with each input, the intended evidence IDs, expected response conditions, expected tool decision, and failure category. For example, a refund-window question should cite the appropriate current sample record; an unsupplied warranty question should not invent a number; a retired policy question should distinguish date and status; and an unauthorized refund command must not trigger a write. Record actual outcomes only when a learner runs the tests in an approved environment.
Prompt-level evaluation is useful but insufficient. If the agent answers correctly because it memorized a policy embedded in system instructions, that does not prove the search integration works. Change the current policy in the test dataset, reindex, and check whether the answer follows the updated source. If it does not, investigate document versioning, retrieval caching, index freshness, and response grounding before tweaking the prompt.
Assess negative examples deliberately. An agent that confidently answers 95 ordinary questions but follows one prompt-injection instruction to reveal protected policy content cannot be considered safe just because its average quality is high. Use explicit pass/fail gates for unauthorized tool use and data exposure; use graded measures for less critical qualities such as readability and explanatory helpfulness. Make abstention an acceptable outcome when the approved evidence is missing.
Observe the full request path and troubleshoot by layer
Instrument one test request end to end: caller identity, conversation ID, model invocation, retrieval event, tool selection, tool outcome, validation decision, and final response. Keep enough correlation data to explain what happened while avoiding raw secrets and sensitive content in general logs. A chat transcript alone may not reveal that retrieval found the wrong document or that a tool was denied for an appropriate security reason.
Track latency by stage. A slow answer might result from authentication, vector search, repeated model reasoning, an external tool, or a retry loop. The production agent telemetry article provides a useful architecture for correlating those parts. It is not enough to measure only the final model response time, because an otherwise fast model may sit behind a slow retrieval or business-service dependency.
Classify failures before retrying. Authentication failures require identity and permission checks; a missing model requires deployment verification; empty retrieval requires indexing and query inspection; a tool argument mismatch requires schema correction; a fabricated refund policy requires grounding and evaluation work; and service throttling requires capacity or backoff changes. Retrying every error as though it were transient can conceal the underlying issue and create duplicate tool actions.
Review traces with the same care as source documents. Observability systems often retain input snippets, retrieved content, exception messages and identifiers. Apply data minimization, redaction and retention rules before enabling broad logging. A technically helpful trace can become a privacy issue if it contains information that should not be visible to the operator reviewing it.
Make the lab independently checkable before release
A finished training exercise needs a manifest of its own: current source document versions, SDK requirements, model deployment, sample inputs, expected evidence, expected denials, evaluator criteria, and cleanup steps. The lab should be reproducible by another authorized learner, but it should not claim that the writing team executed Azure resources when it did not. Include clearly labeled illustrative responses rather than screenshots of imaginary successful cloud runs.
Protect the distinction between a working demonstration and an approved production release. To move toward production, a team would still need cost and scaling estimates, policy and security review, load and resilience testing, model/version lifecycle assessment, ownership of the source corpus, rollback procedures, and change-management approval. A lab can exercise those decisions in miniature without pretending to replace them.
For the current AI-103 role, the important outcome is the ability to explain why every stage exists: why a project uses an authenticated identity, why the agent is limited to a defined task, how its response obtains evidence, why a tool cannot authorize itself, how a test proves that an unsafe action was blocked, and how a trace identifies a failure. Those are skills that carry beyond one SDK version or model name. An honest lab records what was verified, what was simulated, and what remains untested.