Structured outputs in Microsoft Foundry Models make Azure OpenAI responses follow a supplied JSON Schema instead of merely returning syntactically valid JSON. Current Microsoft documentation supports structured outputs through both the Responses API and Chat Completions API: Responses uses text.format, while Chat Completions uses response_format. This is stricter than older JSON mode, which guaranteed valid JSON syntax but did not guarantee schema adherence.
Within Microsoft AI Agents, structured outputs are the contract between probabilistic model behavior and deterministic application code. Azure OpenAI Responses API provides the surrounding API model.
Use structured outputs when downstream code depends on shape
Extraction, classification, routing, workflow state and function/tool planning all benefit from a schema that rejects malformed shapes before they reach application logic.
Prompt-only “return JSON” remains appropriate for human-readable experiments but is a weaker production contract.
Use schema-constrained outputs when a parser or automated action depends on the result.
Responses API and Chat Completions configure schemas differently
Current Microsoft guidance uses text.format for Responses and response_format for Chat Completions.
Keep API-specific request builders centralized so one schema definition can be reused without duplicating business semantics.
Do not mix examples from the two APIs blindly; the payload surface differs.
Keep schemas aligned with the consumer contract
A schema should represent exactly what the next system needs: stable IDs, required fields, enums, nullability and nested objects.
Avoid dumping every model thought or explanation into the JSON contract.
Human-readable rationale can be a separate field only when the consumer actually uses it.
Use enums for closed business states
Statuses such as approve, deny, review can be constrained reliably.
Open-ended categories should not be frozen into an enum simply for convenience.
Schema evolution becomes an application deployment concern, so choose closed sets deliberately.
Represent unknown rather than forcing fabrication
If the source does not contain a value, allow null or an explicit unknown state.
Do not require a string for every field if the model would have to invent information to satisfy the schema.
Structured outputs guarantee shape, not factual truth.
Schema support has limitations
Microsoft documents supported JSON Schema features and limitations for structured outputs.
Validate schemas before deployment and keep them within the supported subset.
Complex recursive or highly dynamic schemas should be tested against the exact deployed model/API version.
Function calling benefits from the same discipline
Structured outputs can be used with function/tool calling so parameters conform to a defined structure.
For write-capable operations, also validate business authorization outside the model.
A schema-valid request can still be unauthorized or unsafe for the current user.
Version schemas with model and prompt changes
Changing required fields or enum values is a breaking API change for consumers.
Version the schema or use an expand-contract rollout.
Azure OpenAI Model Versioning should be coordinated with output-contract changes so regressions can be attributed correctly.
Use deterministic tests for structured correctness
Evaluation can assert parse success, types, required fields and business invariants directly in code.
Reserve LLM judges for semantic quality that cannot be checked deterministically.
This lowers evaluation cost and makes CI failures easier to diagnose.
Watch prompt and token overhead
Schema definitions add request context.
Very large schemas can raise token use and increase maintenance cost.
Keep the contract concise and split unrelated workflows rather than creating one giant “everything” response object.
Azure OpenAI structured outputs succeed when schema and application policy stay separate
The mature design uses schema-constrained output for machine contracts, explicit unknown states, versioned schemas, deterministic tests and application-side authorization.
Structured outputs make responses parseable and type-safe; the application still owns whether the requested operation is valid, permitted and correct.
Schema ownership should sit with the service that consumes the model result. Prompt authors can propose field names, but downstream APIs, databases and workflows define compatibility requirements. Store the schema in source control and generate typed clients or validators from one canonical definition where practical.
Azure/OpenAI model support for structured outputs should be checked in the current supported-model table before deployment. A schema that works on one model/version or API version may not be supported on another. Tie the model deployment and schema test suite together so a model migration cannot bypass compatibility checks.
Function calling and final structured output should use separate contracts. The tool schema describes what the model may ask the application to do; the final response schema describes what the application receives after tool execution. Keeping them separate reduces coupling and allows a workflow to call many tools while still returning one stable result object.
Business validation should happen after schema validation. An integer `quantity` can still be negative; an enum `approved` can still be inappropriate for the user’s role. Validate ranges, cross-field rules, external IDs, permissions and current state in deterministic code before committing side effects.
Use schema descriptions for normalization, but avoid writing a second entire prompt inside the schema. Long descriptions add tokens and can become inconsistent with the system prompt. Keep domain semantics in one place and have concise field-level descriptions for format or ambiguity.
Nullability should distinguish absent from intentionally empty. An optional field omitted by schema, a present `null`, and an empty string can mean different things downstream. Choose one representation and document it so model output maps predictably into databases and APIs.
Structured outputs are especially useful in agent orchestration when the next step is selected by code. Instead of parsing prose to decide ‘search’ versus ‘escalate,’ define a constrained action enum plus validated parameters. This makes the orchestration graph explicit and testable.
Schema migrations should be rolled out with expand-contract patterns. Add a new optional field first, deploy consumers that understand it, then make it required later if needed. Avoid changing enum meaning in place because historical records and cached responses may still use the older semantic.
Evaluation should compare structured correctness and semantic correctness separately. A response can be 100% schema-valid and 70% factually correct. Report parse/schema pass rate, task accuracy, field-level precision/recall, and invalid-business-rule rate as different metrics.
Logging should avoid dumping sensitive structured outputs indiscriminately. JSON is easy to index, which can make PII or secrets more searchable in telemetry than in prose. Apply field-level redaction and retention policies to structured logs just like any other production data.
Schema files should be reusable across languages when multiple services consume the output. Generate Pydantic, TypeScript, C# or Java validation types from one JSON Schema or equivalent source where tooling allows. This reduces the chance that one client interprets nullability or enum values differently.
Structured-output failures should be monitored separately from ordinary model errors. If a deployed model/API version starts rejecting the schema or returns a validation-related error, alert the owning team before application retries amplify load. Include schema version and model deployment in telemetry.
Use a smaller schema for intermediate orchestration decisions than for the final business object. A router may need only `intent`, `confidence`, and `next_action`; the final extractor may need dozens of fields. Keeping intermediate contracts narrow reduces token cost and makes agent control flow easier to audit.
Human review UIs should render structured fields directly rather than dump raw JSON. Reviewers can then edit or approve specific values while the application records which fields changed. This makes human correction part of the data pipeline and produces better evaluation labels.
Structured outputs should be tested against model fallback behavior. If the primary deployment is unavailable and traffic moves to another supported model, run the same schema conformance and semantic tests. Do not assume every fallback has identical structured-output support or quality.
Structured-output API versions should be pinned and tested. Azure model deployments can move across API versions and supported feature sets; a client library update may default to a newer surface. Keep the production API version explicit and migrate only after schema regression tests pass.
Use AI System Change Control to coordinate model upgrades with schema contracts. Model changes and schema changes should normally be separated so a regression can be traced to one cause.
Agent gateways can enforce schema-validation telemetry centrally. API Management for AI Gateways is relevant when several applications share model deployments; a gateway can record model/version, schema version, failures and latency without each service inventing its own logging format.
Schema design should include maximum output size considerations. A valid JSON object with a huge array can still exceed output budgets or downstream storage limits. Define business-level caps and pagination/batch patterns rather than asking one model response to contain an unbounded collection.
Fallback behavior should never silently switch from structured output to ordinary text. If the selected model or API surface cannot honor the schema, fail explicitly or route to a tested compatible deployment. Silent format downgrade is one of the easiest ways to turn a reliable workflow back into brittle parser logic.
Schema contracts should also specify how consumers handle model refusal or safety-filter outcomes. These are not ordinary business objects and should not be coerced into a fake valid result. Route refusal/error states through an explicit envelope or transport-level error path so downstream systems never interpret a blocked generation as an approved structured decision.