The Microsoft Foundry model catalog is not just a list of model names. It is the discovery layer where architecture requirements begin to narrow the model field: provider, region, deployment option, lifecycle stage, inference task, supported features, and model-specific details can all affect whether a candidate belongs in the solution. The useful skill is therefore not “find the biggest model.” It is translating workload requirements into filters and then validating the remaining candidates with model cards, benchmarks, deployment constraints, and your own evaluation.
For a Microsoft AI agent platform, catalog decisions affect everything downstream: latency, cost, feature support, tool behavior, governance, upgrade planning, and the way applications are deployed. The same systems thinking is relevant to Microsoft AI-103. A model belongs in an architecture because its capabilities and operating conditions fit the workload, not because it is the most visible option in the catalog.
Start with workload filters before comparing model names
Foundry’s catalog supports filtering by provider collection, Azure region, deployment option, deployment SKU, lifecycle, industry, supported features, and inference task. Those filters map directly to architectural questions. Does the workload need chat completion or embeddings? Must it support reasoning, tool calling, vision, or another feature? Is the target region available? Is the model in preview, generally available, or moving toward deprecation?
Starting with those constraints avoids a common anti-pattern: selecting a model from reputation and then redesigning the application around its limitations. A workload with strict regional requirements should filter by region first. A production service that cannot accept preview dependencies should remove preview candidates before quality comparison. An agent that depends on structured tool use should screen for the required features before benchmark scores become relevant.
This is the catalog equivalent of foundation-model selection: the decision is a constraint problem before it is a ranking problem. The catalog makes several constraints visible, but the architecture team still has to define them.
Read the model card as an engineering document
A Foundry model card can include quick facts, detailed model information, version and supported data types, existing deployments, benchmark results for selected models, and licensing information. Those sections answer different questions. Quick facts help with screening. The details tab supports capability analysis. Deployment information tells you how the model can be operated. Benchmarks provide standardized evidence. Licensing establishes legal conditions that must be understood before use.
The license tab deserves particular attention for partner and community models. Microsoft classifies models from providers other than Microsoft as Non-Microsoft Products subject to the provider’s terms. That distinction can affect support, procurement, data handling review, and legal approval. A technically strong model that cannot pass the organization’s vendor or license process is not a deployable candidate.
Model cards should be captured in the architecture decision record at the version level. “We use Model X” is insufficient if a newer version behaves differently or the original version enters a lifecycle transition. Record the exact version, important capabilities, deployment choice, evaluation evidence, and the reason it was selected.
Treat lifecycle state as a production requirement
Foundry models move through lifecycle stages. Current Microsoft policy distinguishes stages including Preview, Generally Available, Legacy, Deprecated, and Retired, with different implications for new deployments and continued service. A model can be functionally excellent and still create migration risk if its lifecycle does not fit the expected lifespan of the application.
Preview is appropriate for experimentation when the organization accepts change and reduced guarantees. GA is the normal target for production workloads that require stable support commitments. Legacy or Deprecated states should trigger migration planning rather than surprise. Retired models are no longer available for inference. The catalog’s lifecycle filter therefore belongs in normal model governance, not just in incident response after an endpoint stops accepting traffic.
A mature team maintains a model inventory with owners and review dates. The agent lifecycle and model lifecycle should be linked so that an agent release cannot silently depend on a model whose support horizon conflicts with the application’s operating plan.
Deployment options change the operating model
Foundry currently describes Serverless API as the preferred deployment option and also supports managed compute for appropriate open-source, partner, and custom models. Serverless API supports the broadest range of Foundry Model capabilities and deployment types, while managed compute uses dedicated accelerator capacity managed through the project. Some supported models can also be tried through instant access in preview without creating a deployment first.
These are not interchangeable labels. The deployment option affects where resources live, how capacity is billed, which content filtering capabilities are available, network design, scaling behavior, and the operational work required from the platform team. A catalog decision should therefore include deployment architecture from the beginning rather than choosing a model first and treating deployment as an implementation detail.
For production, compare expected workload economics under the actual deployment type. A model that is attractive at low intermittent volume might not be the best choice at sustained high utilization, and dedicated managed compute can introduce different capacity and operations questions from token-based serverless consumption. Test the combination of model and deployment, because that combination is what the application will run.
Use benchmarks to shortlist, then test on your data
The catalog exposes benchmark metrics and leaderboards for selected models. These can compare quality, safety, estimated cost, throughput, and scenario performance. They are useful for creating a shortlist because they provide standardized evidence across models that would otherwise be difficult to compare consistently.
Public benchmarks still cannot represent every domain. A claims assistant, engineering copilot, or security triage agent has vocabulary, context structure, tool interfaces, and risk boundaries that differ from public test sets. The evaluation and regression process should therefore run shortlisted models against representative workload data before any production decision.
Record case-level results, not just averages. One candidate may be slightly weaker overall but significantly better on the difficult cases that dominate business risk. Another may require more prompt scaffolding or retries, which changes system-level latency and cost. Catalog evidence gets the team to the right experiment; workload evidence makes the final choice.
Feature support should be tested at system level
A catalog feature flag such as tool calling, structured output, or vision indicates model capability, but application reliability depends on how that capability behaves with your schema and orchestration. A model can technically support tool calling yet choose the wrong tool more often in a crowded toolbox. Structured output can be supported yet still require schema design that avoids ambiguous unions or deeply nested structures. Vision support says little about the quality of the specific documents the application must understand.
For agent workloads, test the model with actual tool descriptions, parameter names, conversation history, and failure responses. For retrieval workloads, test the complete context assembly pattern rather than a bare prompt. For structured workloads, validate both schema compliance and semantic correctness. The model is one component in a system, and catalog capabilities need to survive the system around them.
The existing Foundry agent architecture discussion is relevant because model choice interacts with identity, tools, memory, and observability. A model that looks ideal in a playground may behave differently when exposed to the full production capability set.
Plan for model diversity instead of assuming one universal model
The catalog encourages comparison, and many architectures benefit from using more than one model. A high-capability model can handle complex reasoning while a smaller model handles classification, extraction, routing, or simple summarization. An embedding model can be selected independently from the conversational model. Routing can reduce cost and latency when the decision boundary is tested and observable.
Model diversity also reduces migration risk. If the application has a clean abstraction around model invocation and stable evaluation sets, a lifecycle change does not require rebuilding the entire system. New catalog entries can be tested against the existing contract and promoted only when they meet the workload’s thresholds. This turns model churn into a managed engineering process.
Do not create diversity for its own sake. Each additional model adds deployment, evaluation, monitoring, security, and support surface. The right architecture uses the smallest model set that materially improves the workload while keeping behavior understandable.
Build a repeatable catalog-to-production workflow
A practical workflow begins with a written requirement profile: task type, modalities, context needs, tool or structured-output requirements, target regions, deployment constraints, safety needs, latency objectives, throughput, and budget. Apply those requirements as catalog filters. Read the remaining model cards and licensing terms. Use benchmarks to narrow the field. Then evaluate candidates on representative data and under realistic load.
The final decision should produce an artifact containing the selected model version, deployment type, evaluation results, known limitations, lifecycle status, migration trigger, and an owner. Revisit the decision periodically because the catalog and benchmark landscape change. The purpose of the catalog is not to freeze an answer; it is to make discovery and comparison systematic.
When used this way, Foundry’s model catalog becomes an architecture tool. It connects feature discovery, lifecycle governance, benchmark evidence, deployment choices, and licensing into the first stage of model selection. The production decision still belongs to the workload, but the catalog helps teams reach that decision with far less guesswork.