A tool schema is part of the agent’s control surface
Function tools expose a structured contract that tells the model what operation exists and what arguments it may supply. Microsoft Agent Framework can generate schemas from typed function signatures or accept explicit schema definitions when teams need tighter control.
For AI-103, this is not just developer convenience. A schema shapes model behavior, validation, authorization boundaries, observability, and backward compatibility, so changing it can alter the agent even when the underlying business function is unchanged.
Function calling becomes reliable when the tool contract narrows the model’s choices to explicit operations and constrained arguments. The model can select an operation, while deterministic code enforces types, permissions, and business invariants outside natural-language reasoning.
Treat tool definitions like public APIs. Give them stable names, explicit owners, versioning expectations, tests, and a review process for breaking changes.
Tool catalogs should use a naming convention that survives implementation changes. A business-oriented name such as “create_ticket” is more stable for model selection than a backend-specific endpoint name that changes when the service is refactored.
Namespace conventions matter when multiple teams publish tools. Prefixes or domain grouping can prevent collisions without making every visible name excessively long, and discovery metadata can carry ownership details that the model does not need inside the invocation name itself.
Names and descriptions should distinguish intent precisely
Tool names should describe the operation in a way the model can discriminate from neighboring tools. Generic names such as “manage,” “process,” or “update” create ambiguity when several tools could plausibly satisfy the same request.
Descriptions should state what the tool does, important boundaries, and when it should be selected. Avoid marketing language and long examples that bury the decision rule inside prose.
Agent tools are easier to select correctly when overlapping capabilities are separated. If two tools differ only by a subtle permission or resource scope, make that difference visible in both the description and parameter model.
Do not encode authorization as natural-language advice such as “only call for admins.” The description can guide selection, but trusted runtime code must verify the caller’s actual permissions.
Description review should include negative boundaries: what the tool does not do and when a neighboring tool should be used. Clear separation reduces accidental selection when several operations share the same business nouns.
Parameter schemas should constrain what the model can invent
Use types, required fields, enums, ranges, formats, and nested structures to narrow valid input. A free-form string that secretly contains five business fields pushes parsing and validation back into model reasoning, where errors become harder to detect.
Prefer IDs from trusted application context when the user has already selected a resource. Asking the model to regenerate tenant IDs, account IDs, or security scopes from conversation text creates unnecessary opportunities for mistakes.
Make optionality meaningful. If a field changes the operation substantially, consider separate tools instead of one schema with many interacting flags that the model must reason about correctly.
Validate again at runtime even if the framework validates JSON shape. Schema correctness does not prove that a date is allowed, a record exists, or the caller may act on it.
Schema generation from code types is convenient, but internal object models can expose fields that users should never control. Review the generated JSON surface and create explicit request types when the backend structure is broader than the agent contract.
Schema defaults should be used sparingly. A hidden default can make a call succeed while doing something the user never intended, especially for scope, time range, environment, or destructive-mode fields; require explicit values when the choice materially changes the effect.
Separate read tools from effectful write tools
Read operations and state-changing operations have different risk profiles. Splitting them into separate tools makes selection, approval, auditing, and least-privilege credentials easier to reason about.
A “get customer” tool can often run automatically, while “close account” should require stronger identity and confirmation. Combining them into a universal customer tool hides the moment when the workflow crosses from information retrieval into side effect.
Agent access must be enforced around the tool rather than inferred from conversational confidence. The runtime must know who the user is and whether the requested operation is allowed.
For destructive actions, include idempotency or confirmation tokens when the underlying system supports them. Agents can retry after timeouts, and a repeated write must not silently duplicate a real-world effect.
Write tools should return a durable transaction identifier when possible. That identifier lets the agent report what happened, lets operators trace the side effect, and helps retry logic determine whether an apparently timed-out operation actually succeeded.
Return schemas and error shapes guide recovery
Tool outputs should be structured enough that the model can distinguish data from status. A plain paragraph such as “could not update record because access was denied” is harder to handle reliably than a stable result with status, error code, safe message, and relevant identifiers.
Classify failures into categories such as invalid input, permission denied, not found, conflict, transient dependency failure, and business rejection. The agent can then ask for missing information, stop on authorization failures, or retry only conditions that are genuinely transient.
Avoid returning stack traces, raw database errors, or secret-bearing payloads to the model. Diagnostic detail can go to protected logs while the tool returns a bounded error contract appropriate for agent reasoning.
Keep success output focused too. Dumping an entire backend object into the context increases tokens and can expose fields the model did not need to complete the task.
Error contracts should avoid inviting the model to retry permanent failures. Include machine-readable categories and safe retry guidance so authorization denials, validation errors, and conflicts are not treated like transient network timeouts.
A tool should expose whether a failure is safe to retry. That signal can be part of structured metadata or runtime policy, but it should not depend on the model guessing from a prose error message because retries can duplicate effects or increase load.
Schema changes need compatibility and rollout discipline
Renaming a field, tightening an enum, or changing a required parameter can break tool use even when the backend API still works. Agent prompts and learned selection behavior may also depend on names and descriptions, so schema changes should be treated as interface releases.
Use versioned tool definitions when a change is breaking. Keep the old version available long enough for consuming agents to migrate and evaluate the new contract on representative tasks.
Agent orchestration magnifies compatibility risk because one schema can be consumed by coordinators, specialist agents, and shared workflows. A small central change can affect many paths that are not obvious from the tool implementation.
Record which agent version used which tool version in traces. Without that evidence, incidents after a rollout can be difficult to separate from prompt or model changes.
Deprecation should be visible in both documentation and runtime telemetry. Owners need to know which agents still invoke an old schema before they remove it, especially when tool discovery is dynamic and consumers are not listed in a static code dependency graph.
Evaluate selection and arguments, not only final task success
An agent may finish a task after choosing the wrong tool several times, passing unnecessary parameters, or relying on retries. Final success hides inefficiency and can conceal a dangerous pattern that becomes costly at production scale.
Build tests for tool selection, required argument extraction, ambiguous requests, missing information, invalid enums, permission denial, and side-effect confirmation. Include near-neighbor tools so the model has to make the same distinctions it faces in production.
Measure parameter correctness and unnecessary calls alongside task completion. A tool schema is successful when it helps the model reach the correct operation with minimal repair, not merely when enough retries eventually succeed.
Use failed production traces as new tests. Real users reveal ambiguous language and edge cases that schema authors rarely anticipate during initial design.
Evaluation should include adversarial parameter requests such as hidden resource IDs, excessive ranges, wildcard scopes, and conflicting fields. These cases test whether the schema and runtime validation jointly constrain the model’s ability to construct unsafe calls.
Good tool schemas move authority out of natural language
The model should reason about user intent, but deterministic systems should own identity, authorization, validation, side-effect controls, and durable business rules. A strong schema makes that division visible by limiting what the model can propose and giving runtime code clear inputs to enforce.
For Microsoft AI agents, Agent Framework, Foundry tools, MCP servers, and shared toolboxes all rely on contracts that models can discover and invoke. The framework may generate JSON schema automatically, but automatic generation does not remove the need for editorial and security design.
Prefer a small set of clear operations over a huge schema that mirrors every backend capability. The model’s tool surface should represent user-relevant actions, not the internal shape of the enterprise API catalog.
The best tool contract is boring: its name is obvious, parameters are constrained, failures are predictable, permissions are enforced elsewhere, and version changes are visible. That simplicity makes complex agent behavior safer to build on top.
Tool design should be reviewed after production incidents. If operators repeatedly discover that a schema allows ambiguous intent or bundles unrelated effects, splitting the tool can be a stronger fix than adding more prompt instructions around the same unsafe contract.
Contract review should include the human operator experience. When incidents occur, support teams need to interpret tool names, arguments, and outcomes quickly, so schemas that are elegant for code generation but opaque in logs can still be operationally poor.