Separate the durable agent architecture from the retiring workflow surface
A Foundry agent architecture can support agent hosting and multi-agent systems, but one orchestration surface now has an explicit end date: Foundry workflows retire on December 1, 2026. A current Microsoft AI agent design should therefore distinguish the enduring Agent Service capabilities from the declarative workflow feature that is being retired.
This distinction matters for Microsoft AI-103 preparation because the technology decision must reflect current platform direction, not just terminology in an older diagram. Microsoft recommends Microsoft Agent Framework for new workflow development, while existing Foundry workflows have migration paths.
Do not start a new production orchestration by treating the retiring visual workflow builder as the default future-state runtime. First decide whether you need custom code orchestration, hosted agents, Agent Framework, or another business-process surface, then design the agent boundaries around that supported path.
Multi-agent structure should follow task topology
Multi-step agent reasoning justifies multiple agents only when roles need distinct instructions, tools, data boundaries, or context; extra agents are not a quality feature by themselves. That keeps the orchestration topology tied to real work boundaries and makes each specialist easier to test independently.
A coordinator can delegate research, extraction, validation, or domain-specific reasoning to specialist agents and then synthesize the result. Sequential work fits dependencies; parallel branches fit independent subtasks; handoff patterns fit cases where responsibility genuinely changes from one specialist to another.
Keep each role narrow enough that ownership is visible. If every agent can call every tool and make every decision, the system becomes harder to secure and evaluate than a single well-scoped agent.
Map information dependencies before assigning roles. If a compliance reviewer needs the exact artifact produced by a generator, that is a sequential handoff; if three specialists can assess independent risk dimensions from the same source, they can run concurrently. Explicit dependency mapping prevents a visual orchestration from hiding unnecessary serial waits or unsafe parallelism.
Keep role prompts stable and versioned. When the coordinator, specialist instructions, and tool catalog all change at once, a regression is difficult to attribute. Production orchestration benefits from treating each role as a separately testable component with clear inputs and outputs.
State and handoffs need an explicit contract
A handoff should specify what the receiving agent gets: user objective, relevant conversation state, verified facts, artifacts, unresolved questions, and the authority it is allowed to exercise. Copying the entire conversation is easy but can pollute specialist context with irrelevant instructions and increase token cost.
Use durable run identifiers and artifact references so orchestration does not depend on conversational memory alone. The broader Microsoft agent stack includes several execution surfaces, and explicit state makes transitions between them easier to audit and migrate.
Treat partial completion as structured state. A specialist should be able to return findings plus gaps or errors, allowing the coordinator to choose whether to retry, assign another worker, ask the user, or continue with bounded uncertainty.
Use schemas for handoff payloads where the next agent depends on specific facts. Fields such as objective, artifact reference, evidence, confidence, open issues, and allowed actions are easier to validate than a long prose summary whose omissions are hard to detect.
Keep conversation history separate from durable business state. A chat transcript can help the next agent understand intent, but order status, approval state, deployment identity, and other authoritative facts should live in systems designed to maintain those facts under retries and concurrent access.
Identity and tool authorization must survive delegation
Each agent should receive only the tools and credentials its job requires. Delegating a task should not implicitly delegate the coordinator’s full privilege set, because a specialist prompt can be manipulated independently and may not need authority over the same resources.
Preserve user and workload identity across tool calls where the backend requires it, while keeping service credentials isolated from model-visible context. Authorization should be re-evaluated at the execution boundary rather than inherited from a natural-language handoff.
For actions that change external state, centralize policy where practical. Specialists can recommend operations, but a controlled executor can enforce tenant boundaries, resource permissions, approvals, idempotency, and audit logging consistently.
Delegated agents also need clear data visibility. A specialist that only needs a summary should not automatically receive the raw customer record, attachment, or secret-bearing context used by the coordinator. Minimize shared context at the handoff boundary and fetch sensitive detail only when the specialist’s task requires it.
Tool ownership should remain visible after delegation. Record which agent requested the operation and which principal executed it so an audit can distinguish a coordinator decision from a specialist proposal even when both eventually use the same backend service.
Failure recovery is part of orchestration design
Multi-agent systems create more failure points: a worker can time out, return malformed output, lose access to a tool, exceed quota, or produce a result that conflicts with another worker. The coordinator needs typed outcomes and retry policy rather than assuming every branch ends in useful prose.
Retries should preserve task identity and avoid duplicating side effects. If a worker already launched a job or updated a record, the recovery path should inspect that operation before creating a replacement request.
Bound recursive delegation and total fan-out. An agent that can continually create more work can consume budget without improving the answer, so production orchestration needs maximum depth, worker count, time, and cost.
Design compensation where a multi-step workflow spans systems without a shared transaction. If one agent creates a resource and a later step fails, the system needs a defined choice: retain the partial resource and surface it, clean it up safely, or mark the run for operator intervention. Pretending the whole workflow was atomic creates orphaned state.
Classify failures by ownership. A model refusal, tool authorization denial, malformed specialist output, external service timeout, and coordinator logic bug should route to different remediation paths. One generic retry policy can amplify outages or repeatedly ask an agent to solve a problem that only an administrator can fix.
Human review belongs where consequences converge
A multi-agent plan can gather evidence autonomously and still require a person before an irreversible action. Human oversight should place approval where evidence turns into consequence rather than on every harmless internal message between agents.
The reviewer should see the normalized proposed action and its provenance: which agents contributed, what evidence they relied on, what checks ran, and what uncertainty remains. A final recommendation without provenance is difficult to challenge when several autonomous branches shaped it.
Denial should become orchestration state. Feed the reason back to the coordinator so it can revise the plan or stop instead of repeatedly asking for approval on an unchanged action.
Observability must show fan-out, handoff, and synthesis
Multi-agent traces should connect the top-level request to every worker, tool call, artifact, and synthesis step. AI observability is especially important here because a single latency or token total cannot reveal which branch caused cost growth or introduced a faulty fact.
Record task identifiers, agent roles, handoff reasons, tool outcomes, retry counts, model selections, and output sizes. The goal is to reconstruct the decision path without logging sensitive payloads indiscriminately.
Evaluate synthesis as its own stage. Strong specialists can still produce a poor final answer if the coordinator drops minority evidence, merges contradictory claims, or overweights the most verbose worker.
Set service-level objectives for the orchestration path, not just individual agents. Fan-out can make median worker latency look healthy while the overall request waits on the slowest branch, so track critical-path duration, cancelled branches, and time spent in synthesis as separate components.
Migration planning should reduce dependence on a retiring feature
Existing Foundry workflows should be inventoried before the December 1, 2026 retirement date, with each flow classified by orchestration pattern, integrations, state, approvals, and operational dependencies. That inventory determines whether a migration to Agent Framework is straightforward or whether the workflow really belongs in a different automation surface.
Avoid cosmetic migration that reproduces every old node one-for-one. Use the transition to simplify duplicated agents, centralize policy, clarify tool contracts, and remove state that existed only because of the old workflow engine.
The durable goal is not to preserve a particular visual diagram. It is to preserve the business behavior, security boundaries, recoverability, and evaluation evidence while moving orchestration onto a supported architecture that can evolve after the retired surface disappears.
Start migration with observability parity. Before replacing a workflow, capture how the existing system records runs, errors, approvals, tool calls, and business outcomes so the new implementation does not lose the evidence operators rely on during the transition.
Run old and new orchestration against a controlled evaluation set before cutover. Compare not only final answers but also side effects, latency, cost, approval frequency, failure recovery, and tool authorization. A migration is complete when the operational behavior is at least as understandable as the feature it replaces.
Assign an owner and cutover date to each retiring workflow instead of keeping migration as a platform-wide backlog item. Dependencies such as custom tools, approvals, external queues, and human handoffs can then be retired in a controlled sequence, with unresolved blockers visible before the December 1, 2026 deadline becomes an outage risk.