A Model Context Protocol server is not simply a collection of functions exposed to Claude. It becomes an integration boundary through which an agent discovers capabilities, reads data, invokes operations, and potentially reaches systems with real side effects. In Claude Engineering, MCP server design should therefore be approached like public API design: define a narrow capability surface, authenticate correctly, make tool behavior predictable, return structured evidence, and keep failures observable.
The MCP ecosystem now has a stable 2026-07-28 protocol revision reflected in current official SDKs. MCP servers can expose tools, resources, and prompts, negotiate capabilities, use transports such as stdio or Streamable HTTP, and integrate authorization for remote services. Claude’s MCP connector can connect remote servers over HTTPS, and Managed Agents can attach MCP toolsets with per-tool permission policies. The protocol standardizes connection mechanics; it does not remove the need for sound product and security design.
Expose business capabilities rather than raw backend endpoints
A tool named run_sql or call_api gives the model too much implementation freedom and too little semantic guidance. Prefer capabilities such as find_customer_orders, create_support_case, or schedule_maintenance_window that map to business intent and enforce domain rules inside the server.
Agent tool design improves when the model chooses among meaningful actions rather than low-level primitives. The server remains responsible for authorization, validation, and invariants that should never depend on the prompt.
Keep tools, resources, and prompts conceptually separate
MCP distinguishes actions from readable resources and reusable prompts. Use a tool when the client needs to perform an operation or computed query. Use a resource for addressable information that can be read by URI. Use prompts when the server provides reusable instruction templates. Blurring these primitives can produce confusing permission and caching behavior.
The distinction also improves observability. API security benefits when a read is visibly a read and a side-effecting operation is visibly a tool call. Reviewers can then apply stricter policies to actions without slowing every informational lookup.
Design the transport around where the server actually runs
Local development tools often fit stdio because the client launches the server process and communicates over standard input and output. Shared or hosted integrations fit Streamable HTTP because multiple clients connect over the network. Claude’s remote MCP connector expects an HTTPS URL for remote servers, while managed-agent deployments can also use private connectivity patterns supported by the platform.
Transport choice changes the threat model. A local stdio server inherits the security of the host process, while a remote server needs TLS, authentication, rate limiting, and network controls. Autonomous agent security should consider MCP transport as part of the trusted computing boundary rather than an implementation detail.
Use OAuth or another supported credential path without exposing tokens to the model
Remote MCP servers often require authorization. Claude’s MCP connector supports authorization tokens, and Managed Agents can use pre-registered vault credentials for MCP sessions. The model should not need to read or manipulate those tokens. Authentication belongs to the connection layer and tool executor.
Authentication architecture should determine whether the server acts as the user, a service identity, or a delegated application. That identity choice affects downstream authorization and audit. A single server-wide bearer token may be convenient but can erase user-level accountability if the business action needs delegated access.
Keep the tool catalog small enough for accurate selection
Large MCP servers can expose dozens or hundreds of tools. More tools are not automatically more capable because the model must still select the right one. Clear names and descriptions matter, and Claude’s tool-search capabilities can defer loading large catalogs so only relevant definitions enter context.
AI cost and performance applies here because huge tool schemas consume context and can reduce selection quality. Split unrelated domains into separate servers or toolsets where ownership and security boundaries support that structure.
Make tool results structured, bounded, and useful for the next decision
A tool result should return the information Claude needs to continue, not a raw database dump or entire API response. Prefer a compact structured result with stable identifiers, status, and important fields. If the operation creates a large artifact, return a reference or resource URI instead of flooding the context window.
This also reduces data leakage. GenAI observability should capture result size, error class, latency, and identifiers without duplicating sensitive payloads. A result that is easy for the model to use is usually easier for operators to debug as well.
Separate read-only and side-effecting tools for permission policy
Claude Managed Agents can apply permission policies to MCP toolsets and override individual tools. MCP toolsets default to asking for approval. That makes it valuable to expose side effects as clearly distinct tools rather than combining read and write behavior under one ambiguous operation.
Approval boundaries should map to the MCP surface. A lookup tool may run automatically, while a tool that deletes, deploys, sends, or purchases should require stronger confirmation. Tool names and descriptions should make that consequence obvious before execution.
Return explicit errors and preserve idempotency for retried calls
Networked MCP calls can fail, time out, or be retried. The server should distinguish validation errors, authorization failures, not-found conditions, transient dependency failures, and internal errors. Side-effecting tools should accept idempotency identifiers where duplicate execution would be harmful.
Agent lifecycle management benefits from stable error contracts because prompts and retry logic can change without rewriting the backend. A tool that reports every problem as an opaque string forces the model to guess whether retrying is safe.
Version and test the server like a production API
Changing a tool name, input schema, output shape, or authorization behavior can break clients even when the server still starts successfully. Maintain contract tests, run the MCP Inspector or equivalent integration checks, and stage incompatible changes. Monitor which tools are actually used before removing a legacy capability.
Anthropic makes MCP a convenient way to connect Claude to external systems, but the protocol should not hide engineering discipline. Expose narrow business capabilities, separate primitives, choose the right transport, keep credentials outside model context, control permissions by effect, return bounded results, and maintain stable contracts. A well-designed MCP server makes Claude more capable without making the surrounding system less understandable.
Multi-tenant servers need especially careful context isolation. A tool request should carry enough authenticated tenancy information for the server to restrict queries and side effects before touching downstream systems. Do not let the model supply an arbitrary tenant ID and assume that makes the call authorized. The server should derive or validate tenancy from the authenticated principal and reject cross-tenant access even if the schema is valid.
Rate limiting belongs at the MCP boundary too. A model can call tools in parallel, repeat failed calls, or generate bursts during a complex task. Per-user, per-tool, and downstream-aware limits prevent the server from becoming an amplification layer against its own dependencies. Return clear throttling errors so the client can back off rather than retrying aggressively.
Health checks should exercise more than the HTTP listener. A server can accept connections while its OAuth provider, database, or downstream API is unavailable. Use synthetic calls for critical paths and expose metrics for connection failures, authentication failures, tool latency, error codes, and result size. Operations teams need to distinguish an MCP transport problem from a backend dependency problem quickly.
Finally, keep server instructions minimal. Tool descriptions and schemas should communicate capability, while security policy, business rules, and authorization live in code and configuration. Overloading descriptions with behavioral policy makes contracts harder to review and can create false confidence that the model will enforce rules the server itself ignores. The server should remain safe even if Claude misunderstands the description.
Data contracts for resources deserve the same attention as tools. Stable resource URIs, content types, pagination, and change notifications make large knowledge surfaces easier for clients to consume without repeatedly calling bespoke tools. If information is naturally read-only and addressable, modeling it as a resource can reduce side effects and make caching behavior clearer.
Server ownership should be explicit. Each MCP service needs a team responsible for schema changes, security patches, credential rotation, availability, and incident response. A large enterprise will otherwise accumulate abandoned servers that remain connectable long after their original project ends. Treat the MCP catalog as part of the service inventory and apply the same lifecycle controls used for other production APIs.
Schema discovery should be fast enough that the agent is not punished for using the correct server. Keep startup and capability negotiation lightweight, cache stable metadata where appropriate, and avoid performing expensive backend calls simply to list tools or resources. If the catalog is dynamic, make changes observable and versioned so clients do not see a different capability surface with no explanation. Predictable discovery is part of server reliability because tool selection begins before the first business call.
Security reviews should include dependency and supply-chain risk. MCP servers often wrap SDKs, databases, SaaS clients, and authentication libraries. Pin and patch those dependencies, restrict outbound network access to required services, and avoid giving the server filesystem or cloud permissions it does not need. A narrow tool contract is valuable only if the process implementing it is equally constrained.