Model Context Protocol servers give Microsoft agents a standardized way to discover and invoke external tools, but the protocol does not remove the need for ordinary application security. An MCP server can expose read operations, write operations, administrative actions, and business workflows through one tool surface. In Microsoft AI Agents, the practical question is therefore not whether an agent can connect to MCP. It is whether the organization can prove which server it trusted, which tools were exposed, which identity was used, what data crossed the boundary, and what happened when the server behaved unexpectedly.
Current Microsoft Foundry Agent Service supports remote MCP servers through project connections and multiple authentication patterns, including project managed identity, agent identity, OAuth-oriented flows, custom keys, and unauthenticated access when a server is intentionally public. That flexibility is useful, but it creates architecture choices that should be resolved before production. A server that is safe for documentation lookup is not automatically safe for financial updates, ticket closure, or infrastructure changes.
Treat the MCP server as a remote execution boundary
An MCP connection is not just a content source. Tool descriptions tell the model what actions are available, schemas define arguments, and tool results feed new information back into the model. That means a compromised or poorly designed server can influence both action selection and subsequent reasoning. API security fundamentals still apply: authenticate the caller, authorize the action, validate inputs, constrain outputs, and assume that a remote dependency can fail in ways the model cannot safely interpret on its own.
Teams should document the server owner, hosting location, data-retention expectations, change process, and incident contact just as they would for any other privileged API. Microsoft guidance specifically recommends tracking the MCP servers added to Foundry and favoring servers operated by trusted providers rather than opaque proxies. The protocol makes integration easier; it does not establish trust for you.
Use an explicit allow list instead of exposing every discovered tool
MCP servers can advertise many tools, some of which may be irrelevant or too powerful for a particular agent. Current Foundry guidance recommends constraining available tools with an allow list. That is a strong default because fewer tools reduce accidental selection, shrink the prompt surface the model must reason over, and make authorization reviews easier. An agent that only needs to search documentation should not receive deployment or deletion tools simply because the same server offers them.
The same principle appears in agent tools and multi-step reasoning: capability should follow task need. Tool minimization is stronger than a prompt that merely says “do not use dangerous operations.” The model should not be asked to police a capability that the architecture could simply withhold.
Choose the authentication mode according to whose action the server should see
Foundry tool connections now support several identity choices because MCP use cases differ. A project managed identity is appropriate for shared service-to-service access when all agents in a project can legitimately operate under one principal. Agent identity is useful when a published agent needs its own least-privilege roles and audit trail. OAuth or user Entra token patterns are more suitable when the downstream service must preserve user-specific permissions rather than accept the application’s shared authority.
This is a business decision as much as a technical one. Identity architecture should answer who is accountable for the operation. If an agent closes a case, changes a record, or queries sensitive data, downstream logs should identify the principal that policy expects. Do not use a shared key simply because it is the quickest authentication option if the action really requires per-agent or per-user attribution.
Keep credentials in connections rather than exposing them to the model
Foundry connections can hold authentication material so agent code and model context do not need to carry API keys or authorization headers. This is an important separation. The model should decide which tool to invoke and with what business arguments; the platform should attach credentials out of band. If a tool requires a custom header, configure it in the connection rather than teaching the model to reproduce a secret string.
For Azure-hosted tools, managed identity can remove the secret entirely. Azure Key Vault secrets remain useful when a remote system still requires a key, but retrieval and injection should happen through trusted application infrastructure. A prompt, chat history, or MCP tool result should never become a credential transport mechanism.
Assume tool descriptions and results are untrusted input
An MCP tool description can influence the model before any invocation occurs, and a tool result can contain text that looks like new instructions. Microsoft guidance explicitly warns that descriptions, annotations, and results from remote MCP servers should be treated as untrusted input. That matters because prompt injection can arrive indirectly through a database field, issue description, webpage, or document returned by a tool.
Data protection boundaries for Copilot illustrate the broader rule: retrieved content must not silently become authority. System instructions, application policy, tool permissions, and approval gates should remain distinct from data returned by the server. Sensitive actions should be validated against structured state, not authorized because a text result told the model that the action was safe.
Design approval points around side effects, not around protocol calls
Not every MCP invocation needs human confirmation. Read-only metadata lookup can often run automatically, while irreversible or high-impact operations deserve stronger controls. The useful boundary is the business side effect: sending a message externally, changing access, spending money, deleting data, publishing code, or altering production state. A generic “approve every MCP call” policy creates fatigue, while no approval policy gives the model too much autonomy.
Agent access and approval should therefore classify tools by consequence. Some operations can be allowed outright, some can require confirmation only above a threshold, and others can be unavailable to the model entirely. The MCP layer should expose enough structured information for the application to enforce those rules before the request reaches the downstream service.
Trace server, tool, identity, arguments, and outcome together
When an agent uses several tools in one workflow, a generic application log saying “MCP call succeeded” is not enough. Production tracing should capture the server identity, tool name, authenticated principal, correlation ID, sanitized arguments, latency, status, and a safe summary of the result. That data allows operators to distinguish a reasoning error from a permission error, a server timeout, or a schema mismatch.
Agent analytics and monitoring becomes especially important when a remote server is owned by another team. A distributed trace should show the model decision, tool call, and downstream response without leaking credentials or sensitive payloads. Retention should match the risk of the data being logged; “more telemetry” is not automatically better if the trace captures private tool content.
Test schema changes and server failures as first-class release scenarios
MCP reduces integration boilerplate, but servers still evolve. A renamed tool, changed argument schema, new required field, or altered result structure can break an agent even when the endpoint remains reachable. Treat server contracts as versioned dependencies. Before a server update reaches production, regression tests should exercise representative tool calls and confirm that the agent still selects the intended operation with valid arguments.
Failure behavior matters too. Timeouts, partial responses, 429s, authentication failures, and malformed structured output should lead to bounded retry or graceful failure rather than improvised model behavior. Reliable LLM chains should include remote-tool degradation because an agent that can reason well but cannot distinguish a failed write from a successful write is not operationally reliable.
Govern the server catalog as part of the agent portfolio
As teams adopt MCP, the number of servers can grow faster than the number of agents. Without governance, different groups connect duplicate servers, use inconsistent authentication, and expose overlapping tools with different trust assumptions. A controlled catalog should record ownership, approved use cases, supported authentication modes, data classifications, allowed tools, and review dates. Retire servers that no longer have an owner or a clear business purpose.
Enterprise agent governance is incomplete if it catalogs agents but ignores the tool ecosystem behind them. Microsoft Foundry provides the connection and identity mechanisms, but the organization still decides which remote capabilities belong inside its trust boundary. A production MCP strategy succeeds when the protocol becomes a governed interface: trusted servers, explicit tool allow lists, correct identities, protected credentials, testable schemas, and observable side effects.
Server discovery also deserves control. An enterprise should not let agents dynamically connect to arbitrary MCP endpoints supplied in user content, because that turns a convenience feature into an open egress and code-execution channel. Keep approved endpoint addresses in deployment configuration, validate TLS and host identity, and make server additions an administrative action. If a server is moved or replaced, treat the endpoint change like any other external dependency change: review ownership, authentication, tool inventory, and data handling before agents are allowed to use it.
A useful preproduction exercise is to run the same business task with the MCP server unavailable, slow, and returning a schema-valid business error. The agent should distinguish “tool not reachable,” “tool rejected the request,” and “tool succeeded with a negative result.” Collapsing all three into one natural-language failure encourages retries that may duplicate side effects or hide policy denials. Explicit failure semantics make remote tools easier to automate safely and easier for users to understand.