A tool schema is the contract between an AI agent and the system it is allowed to call. Good schemas make the model’s decision space smaller and the backend’s validation job clearer; poor schemas force the model to infer hidden business rules from vague descriptions and then ask production APIs to tolerate mistakes. In Microsoft AI Agents, Microsoft Foundry function tools use names, descriptions, and JSON Schema parameters to tell the model what operations exist and what arguments they accept. The engineering challenge is to make that machine-facing interface precise without pretending the model itself is an authorization boundary.
Microsoft’s current function-calling guidance emphasizes clear function descriptions, correct JSON types, required properties, and strict schema validation where supported. Structured outputs can enforce JSON Schema more strongly than older JSON-only modes, and Foundry agents can return multiple function calls in one response. Those capabilities improve reliability, but the application still has to execute tools, validate results, handle errors, and return tool output to the model so the conversation can continue.
Give every tool one clear responsibility
A tool named manage_customer that can search, edit, refund, suspend, and message a customer creates unnecessary ambiguity. Split capabilities around stable business actions such as get_customer_profile, create_refund_request, or send_case_update. Narrow tools make it easier for the model to select correctly, for the backend to authorize specifically, and for operators to understand what a tool call means in traces.
Agent tools and multi-step reasoning become more reliable when the agent composes several explicit capabilities instead of invoking one oversized function with dozens of modes. Tool boundaries should reflect real operational boundaries, not simply mirror whichever internal API happened to exist first.
Write descriptions that explain intent and constraints, not implementation trivia
The tool description should tell the model when the tool is appropriate, what it does, and any important preconditions. Avoid internal service names that mean nothing to the model, and avoid vague descriptions such as “handles orders.” Parameter descriptions should state units, formats, accepted identifiers, and semantic meaning. Examples can help when a field is commonly confused with another field, but they should not replace validation.
API security remains responsible for enforcing constraints. A description saying “only use this for the current user” is guidance to the model, not access control. The tool executor must derive or validate user identity server-side and reject requests that fall outside the caller’s authority.
Use JSON Schema to constrain the model’s argument space
Types, required arrays, enums, minimums, maximums, and object structure can eliminate entire classes of malformed calls. Use strings for identifiers that may contain leading zeros, enums for small closed choices, and nested objects only when the business concept is genuinely nested. Make optional fields truly optional rather than requiring the model to invent empty values to satisfy the schema.
Schema evolution offers a useful analogy: changing a tool contract can break consumers even when the backend still accepts old requests. Version important changes, preserve compatibility where practical, and test older agent prompts against the new schema before deployment.
Names should remain stable once the tool is widely deployed because prompts, evaluations, and telemetry often reference them. If a function must be replaced, introduce the new tool explicitly and retire the old one through a controlled migration. Quietly changing meaning behind the same name makes historical traces misleading and can cause a model to apply learned calling patterns to a different operation.
Use strict structured outputs where correctness matters
Microsoft’s structured-output support can enforce a supplied JSON Schema for function calling, reducing the gap between “valid JSON” and “valid business shape.” Strictness is especially useful when downstream code assumes fields are present or when an invalid value would trigger expensive recovery logic. It should be paired with backend validation because a structurally valid request can still be unauthorized, stale, or semantically wrong.
Generative AI evaluation pipelines should include schema-adherence and semantic-correctness cases separately. One test can verify that calls conform to the contract; another should verify that the agent chose the right tool and supplied values that match the user’s intent.
Design write tools for confirmation, idempotency, and safe retries
Read operations and write operations should not share identical risk assumptions. For consequential writes, include fields that let the backend detect duplicate attempts, validate expected state, or require a confirmation token. Network failures can leave the agent uncertain whether an action succeeded, so a retry must not create a second order, payment, ticket, or deletion.
Agent access and approval boundaries should be reflected in tool design. An approval step is stronger when the approved action can be represented as a specific immutable request that the executor validates, rather than a vague conversational “yes” followed by a newly generated set of parameters.
Return structured errors the agent can reason about safely
Tool failures should distinguish invalid input, authorization denial, not found, conflict, throttling, transient dependency failure, and permanent business rejection. Do not return stack traces or database errors to the model and hope it interprets them correctly. A compact error code plus safe explanation and retry guidance gives the agent enough information to choose the next step.
Agent analytics and monitoring should record these outcomes as dimensions. If a tool has a high schema-validation failure rate, improve the schema or description. If authorization denials spike after an agent update, investigate tool selection or identity propagation rather than treating every failure as a backend reliability issue.
Handle multiple and parallel tool calls deliberately
Foundry agents can produce several function calls in one response when the model believes multiple actions are needed. Parallel execution can reduce latency for independent reads, but it is unsafe when calls depend on each other or create side effects that must be ordered. The application should understand dependency and conflict rather than blindly executing every returned call concurrently.
Agentic orchestration provides the broader context: tools are steps in a workflow, not isolated API requests. If one call generates an identifier required by another, execute them sequentially. If two independent lookups can run together, parallelism may improve responsiveness without changing semantics.
Keep secrets and hidden policy out of tool descriptions
Tool schemas are often visible to the model and can appear in traces or debugging interfaces. Do not place API keys, internal credentials, sensitive business rules, or confidential endpoint details in descriptions. Use server-side connections and managed identity where possible, and reference abstract resource names rather than embedding secrets in prompt-visible metadata.
Autonomous agent security depends on keeping authority outside the model context. The agent should know what a tool is allowed to accomplish, not how to bypass the controls around it. This makes prompt injection less valuable to an attacker because the hidden security boundary is not encoded as instructions the model can be persuaded to ignore.
When a schema changes, test that the current agent can call it and that the executor handles the current model’s behavior. Track tool version, agent version, and schema version in telemetry. For widely reused tools, publish migration guidance and deprecation timelines so one team’s improvement does not silently break several agents.
Agent lifecycle management should promote tool contracts through the same development, test, and production path as prompts and models. Contract testing can verify representative valid calls, invalid calls, authorization failures, and tool-result handling before the agent is allowed to use the new version.
Keep a machine-readable schema and a human-readable contract together. Engineers need to understand side effects, authorization expectations, latency, and retry behavior that JSON Schema cannot fully express. The model needs concise field descriptions; reviewers need the broader operating semantics. Maintaining both prevents documentation from drifting away from the actual callable interface.
Treat the schema as a reliability boundary, not a magic safety layer
A precise schema helps the model form better requests and makes validation more deterministic, but it cannot decide whether a user is entitled to perform the action or whether the action is sensible in the current business state. Keep identity, authorization, business validation, concurrency control, and audit logging in the executor and downstream services.
Microsoft Foundry provides function calling and structured-output capabilities that can make agent tools much more dependable. The strongest designs use those features to narrow ambiguity while preserving conventional software controls around every side effect. Tool schemas should make the safe path easy for the model and the unsafe path impossible for the backend.
Schema design should also account for localization and formatting. Dates, currencies, phone numbers, measurements, and user-entered names can be represented differently across regions. Prefer canonical machine formats in the tool contract and let the conversational layer handle presentation. This reduces ambiguity and prevents a locale-specific string from being interpreted differently by the model and the backend.
For tools that accept free-form text, keep that freedom narrow. A short reason or note field may be appropriate, but entire SQL statements, shell commands, or arbitrary URLs dramatically expand the execution surface. When flexible input is unavoidable, route it through a dedicated validation or sandbox layer rather than treating the JSON schema as sufficient protection.