Agent diagrams are easy to draw. A box for a model, a box for tools, a box for knowledge, and arrows between them can make an architecture look complete long before the difficult decisions have been made. A useful design review challenges what the diagram hides: identity, state, failure domains, tool side effects, retrieval quality, model choice, observability, and the ownership of decisions when the system behaves unexpectedly.
Those questions align with the current AI-103 path. Microsoft describes the role as building, managing, and deploying agents and AI solutions with Microsoft Foundry. The platform name itself has evolved from Azure AI Studio and Azure AI Foundry to Microsoft Foundry, but the architecture problem is durable: compose models, agents, tools, data, and Azure controls without losing the ability to reason about the system.
The most effective review starts with a business transaction and follows it end to end. That exposes where state changes, where identities cross boundaries, where retries can duplicate work, and where a component can fail independently. Architecture quality becomes visible through behavior rather than through the number of services on the page.
Start with the outcome and the authority the agent actually needs
An agent that answers product questions, an agent that updates service records, and an agent that changes infrastructure may use the same underlying model family, but they should not share the same authority. Define the outcome, the permitted side effects, and the maximum acceptable consequence before selecting tools. This makes least privilege a design input rather than a cleanup task.
The identity architecture around the solution should distinguish user identity, application identity, agent identity, and downstream resource permissions. A reviewer should be able to answer which principal performed every sensitive action and whether that principal had only the rights required for that action.
Choose an agent type based on control, not novelty
Microsoft Foundry supports managed prompt-style agents and more flexible hosted approaches. The choice changes how much infrastructure, code, scaling, and runtime behavior the platform manages for the team. A simple agent with standard tools may benefit from a managed configuration model, while a custom runtime may be justified when the organization needs specialized frameworks, libraries, networking, or execution logic.
The design review should ask what custom control is actually needed. Bringing more code can increase flexibility but also increases patching, deployment, observability, and operational responsibility. Managed features reduce those burdens but may constrain implementation choices. The right answer follows the workload, not the desire to use the most sophisticated agent pattern.
Treat tool boundaries as transaction boundaries
Tools convert model intent into real effects. Every tool should therefore have a narrow purpose, typed inputs, explicit authentication, validation, timeouts, retry semantics, and a defined failure contract. A generic “call backend API” tool is harder to govern than a small set of actions whose consequences are easy to understand.
The distinction becomes critical when a tool is not idempotent. If an agent retries after a network timeout, the system must know whether the first call succeeded. Payment, provisioning, messaging, and record-creation operations need transaction identifiers or other safeguards. Model reasoning cannot compensate for an ambiguous transaction state.
Grounding quality is an architecture dependency, not a prompt detail
An agent grounded in enterprise data inherits the quality, freshness, permissions, and structure of that data. Retrieval architecture must define authoritative sources, chunking or indexing approach, metadata filters, authorization trimming, and how stale content is removed. A powerful model cannot reliably answer from information that retrieval never returns.
The general foundation-model mental model helps separate model capability from application context. The model supplies language and reasoning capability; the surrounding system supplies trusted business evidence. Architecture should make that separation visible so teams know which layer to investigate when an answer is wrong.
State should be explicit enough to survive retries and handoffs
Agents often maintain conversation context, task state, tool results, and long-running workflow status. If that state lives only in an in-memory conversation, recovery becomes fragile. The system should define what must persist, how long it persists, which identifiers connect a resumed interaction to prior work, and which data should deliberately not be retained.
Multi-agent designs make this harder because state crosses responsibility boundaries. A handoff needs a contract: what context is transferred, what is authoritative, what can be recomputed, and what permissions apply to the receiving agent. Without that contract, adding agents can create hidden coupling instead of modularity.
Scale changes behavior through quotas, latency, and contention
A prototype with one user can hide rate limits, model quota constraints, tool bottlenecks, search latency, and downstream concurrency limits. Scale testing should therefore model bursts, long conversations, repeated tool calls, and the failure of shared dependencies. An architecture that scales the agent endpoint but overwhelms a single database or external API is not scalable.
Cost is part of scale as well. More retrieval, longer context, additional agents, and retries all increase work per transaction. Architects should measure latency and cost by completed task rather than by isolated model call. This exposes designs where orchestration complexity is consuming more resources without improving the outcome.
Design graceful degradation for partial failure
Not every dependency needs to fail the whole experience. If a recommendation tool is unavailable, the agent may still answer from knowledge. If retrieval is unavailable, the system may need to refuse questions that require proprietary facts rather than silently answer from model memory. If a write action fails, the agent should report that the transaction did not complete and preserve enough state for recovery.
The review should classify dependencies as required, optional, or deferrable. Then it should test those failures deliberately. This converts resilience from “the platform is highly available” into a precise statement about what the user experiences when a component is unhealthy.
Observability must connect model behavior to system behavior
Logs from individual services are not enough if the team cannot follow one business request across the agent, retrieval, model, tool, and downstream system. Correlation identifiers, traces, evaluation signals, tool outcomes, and policy events should make the path reconstructable. Sensitive prompt or data content may require masking or controlled retention, but absence of telemetry is not a privacy strategy.
The operational relationship with AI-300 is important because production AI requires monitoring, evaluation, and lifecycle control. AI-103 focuses the builder on implementation, while AI-300 provides a natural adjacent destination when the discussion moves into operationalization at scale.
A defensible design has clear reasons for every boundary
Challenge the architecture with five questions: what authority does this component need, what state does it own, how does it fail, how is its behavior observed, and who operates it? If any box has vague answers, the design is not finished. The Microsoft platform offers many ways to compose an agent solution, but a larger service inventory is not evidence of better architecture.
The strongest Foundry design is the smallest system that meets the outcome while keeping identity, state, failure, and evidence understandable. That principle remains useful as APIs and product names change. It turns the architecture review from a tour of features into a disciplined examination of behavior under real constraints.
Network boundaries deserve the same scrutiny as identity. A model endpoint may be public while data stores use private endpoints, or the reverse. Tools may call software-as-a-service endpoints outside the Azure virtual network. Architects should draw the actual egress and ingress paths, identify which connections rely on public internet routing, and decide whether private networking is required for the data classification involved. A secure identity does not automatically make an unrestricted network path acceptable.
The design should also define its change boundary. Models, prompts, tools, indexes, and agent versions evolve at different speeds, and a production incident may involve only one of them. If every change requires redeploying the entire stack, teams may avoid necessary updates; if every component changes independently, compatibility can drift. A release manifest that records the compatible versions of the important components provides a practical middle ground and supports reproducible rollback.
Data residency and dependency location can also change the answer. A model may be available in one region while a required search or tool service is constrained to another, creating cross-region data movement, latency, or compliance questions. The architecture review should map where prompts, retrieved content, tool payloads, and logs are processed and stored. Region choice is therefore a data-governance decision as well as a capacity decision.
A review should also ask how the architecture is tested outside the happy path. Inject tool timeouts, retrieval misses, permission denials, malformed responses, quota pressure, and agent handoff failures. Then observe whether the system degrades predictably and whether operators can identify the failed layer from telemetry. Resilience is not the absence of errors; it is the ability to contain an error, preserve state, and recover without creating a second failure elsewhere in the workflow.
Capacity planning belongs in the same review. Shared quotas across model deployments or projects can create noisy-neighbor effects that are invisible in a component diagram. Define expected concurrency, burst behavior, quota ownership, and the response when capacity is exhausted. A controlled 429 with retry or queue behavior is better architecture than an unexplained timeout that encourages clients to retry aggressively.