Tool Use and Function Calling: Context Before Defaults

Function calling gives a model a structured way to request work from software outside the model. The model does not execute the business operation by itself; it selects a named capability and proposes arguments, while application code validates the request, invokes the real system, and returns the result. That separation is simple on paper, but production reliability depends on treating the tool boundary as an API contract rather than as a conversational convenience.

Tool integration is a core part of the current AI-103 agent scope. Microsoft Foundry agents can use custom functions and other tools to reach real-time data or actions. The design decision is not “should the agent have tools?” but which tools belong in the agent, how narrowly they should be defined, what identity they use, and what the application does when the model requests something invalid, ambiguous, or risky.

A good default is to keep tool semantics deterministic even when tool selection is probabilistic. The model may choose when a function is appropriate, but validation, authorization, idempotency, and transaction integrity should remain conventional software responsibilities.

A tool schema is part of the prompt surface

The function name, description, parameter names, types, enums, and required fields all influence whether the model chooses and populates a tool correctly. Vague schemas force the model to infer business rules that should have been explicit. Overloaded tools with many optional modes make selection harder and increase the chance of invalid combinations.

Design tool contracts for one clear responsibility. If two actions have different consequences or authorization rules, separate them even if the backend API combines them. A narrow contract improves model selection, makes validation easier, and produces more meaningful logs when something goes wrong.

Tool selection should be tested with negative cases

It is not enough to prove that the model calls a tool when it should. Test that it does not call the tool when the user is only asking for explanation, when required facts are missing, when the requested action is outside policy, or when a different tool is more appropriate. Negative cases reveal whether the description and orchestration boundary are actually understood.

The adjacent AB-620 builder context is relevant because tool-connected agents operate inside business workflows. A good evaluation set includes realistic ambiguity: “show me the invoice” versus “refund the invoice,” or “what would this change do?” versus “make this change.” The words may be similar while the authority required is very different.

Validate every argument as if it came from an untrusted client

The model can produce syntactically valid but semantically dangerous arguments. A date may fall outside the allowed period, an amount may exceed policy, a resource identifier may belong to another tenant, or a free-text field may contain content that should never reach the backend. Schema validation is necessary but not sufficient.

Business validation should run in application code or the destination service. Range checks, ownership checks, policy rules, data classification, and entitlement decisions need deterministic enforcement. The agent can explain why a request was rejected, but it should not be able to bypass the rule by rephrasing the argument.

Identity should be chosen per action path

Some tools should act as the signed-in user, preserving delegated authorization. Others should use a workload or agent identity because the task is asynchronous or system-owned. The choice changes auditability and privilege. A powerful service identity can make development convenient while silently bypassing the access model the business expected.

This is where the Entra identity mental model becomes practical. For every function, record which principal authenticates, which resource authorizes it, and whether the user’s own permissions are supposed to constrain the result. Tool design and identity design cannot be separated.

Side effects need idempotency and transaction awareness

Reads are usually easier than writes. A tool that creates an order, sends a message, changes a configuration, or transfers value needs protection against duplicate execution. Networks fail, clients retry, and agent workflows can re-enter a step after partial errors. Without an idempotency key or equivalent business guard, a harmless retry can create a second real-world effect.

The tool response should make transaction state explicit. “Request accepted,” “completed,” “failed before commit,” and “unknown outcome” are different states. If the backend cannot distinguish them, the agent should not invent certainty. It may need to query status before deciding whether retry is safe.

Timeouts and long-running work change the conversation model

Not every tool returns quickly. A data export, provisioning job, or external approval may outlive the conversational request. The architecture should decide whether the agent waits, polls, receives a callback, creates a task, or tells the user that work is pending. Pretending that every function is synchronous creates brittle experiences and encourages uncontrolled retries.

Long-running operations also need durable state. The conversation should be able to resume after a client disconnect without losing the transaction identifier or approval status. That state belongs in an application or workflow store, not only in model context.

Human approval belongs around consequence, not around every tool

High-consequence functions may require confirmation or approval before execution. The approval should display the proposed action in deterministic terms: target, amount, resource, recipient, and important parameters. Asking a user to approve a paragraph of model-generated prose is weaker because the actual side effect can remain hidden inside it.

Low-risk tools should not be burdened with unnecessary approvals. Read-only retrieval, formatting, or reversible drafts can often proceed automatically. Risk-tiering tools keeps the experience usable while preserving human control where the cost of a mistake is high.

Observability should record intent, request, result, and policy outcome

Tool telemetry needs enough structure to answer what the agent intended, which function it selected, the validated arguments, which identity executed it, what the destination returned, and whether any policy blocked or modified the action. Sensitive values may need masking, but removing all detail makes incident analysis impossible.

The operational discipline associated with AI-300 is useful once these agents reach production. Teams should track tool-call success, validation failures, retries, latency, exception categories, and human overrides. Changes in those rates can expose a prompt regression, backend change, or new user behavior before it becomes a major incident.

The right tool boundary makes the agent easier to trust

Consider an agent that helps employees book business travel. One tool searches approved options, another creates a tentative itinerary, and a third commits a purchase after policy validation and user confirmation. Separating those capabilities preserves different levels of authority and gives the agent safer intermediate states. A single “manage travel” function would be harder to reason about and harder to audit.

The broader enterprise assistant idea becomes valuable when the agent can act, but action makes software engineering more important, not less. The durable rule is simple: let the model choose among well-designed capabilities, while deterministic systems continue to enforce the contracts that protect data, money, and infrastructure.

Tool discoverability can itself become a scaling problem. An agent with dozens of overlapping functions may spend more reasoning effort selecting among them and may confuse tools whose descriptions differ only slightly. Grouping capabilities by domain, exposing only contextually relevant tools, or routing to a specialized agent can reduce that ambiguity. The architecture should measure wrong-tool selection, not just failed calls, because an incorrect function can succeed technically while producing the wrong business outcome.

Output validation is equally important. A backend may return a successful response that still contains values the model should not expose, or a function may return unstructured text when the next step expects a stable schema. Normalize tool responses before returning them to the model, remove unnecessary sensitive fields, and distinguish user-facing data from diagnostic metadata. This keeps a backend implementation detail from becoming accidental prompt context.

Tool governance should include retirement. Functions change, APIs are deprecated, and business rules evolve. Removing an old tool abruptly can break stored workflows, while leaving it indefinitely can preserve unsafe behavior. Version tool contracts where compatibility matters, monitor actual usage, and give the orchestration layer an explicit migration path. A tool catalog is healthiest when it records ownership and lifecycle, not merely availability.

A production catalog should also record whether a tool is safe to call speculatively. Read-only lookups can sometimes be parallelized, while write operations, expensive searches, and rate-limited partner APIs should usually be invoked only after the agent has enough evidence. This distinction lets orchestration optimize latency without turning every possible tool into background work that the user never requested.

Error messages returned to the model should be designed as carefully as success payloads. A backend stack trace is noisy and may disclose implementation details, while a generic ‘failed’ response gives the agent too little information to recover. Return structured error categories such as validation_failed, not_authorized, dependency_unavailable, or retry_later, along with safe fields the orchestration layer can act on. This lets the agent choose a sensible next step without learning secrets it never needed.

Security review should include abuse through legitimate tools. An attacker may not need to inject code if they can persuade the agent to call a permitted function repeatedly, enumerate data, or combine several harmless operations into a harmful sequence. Rate limits, purpose checks, anomaly detection, and narrow scopes help contain that class of misuse.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!