Running AI in production on Azure is an architecture problem before it is a model problem. The current AI-300 study guide explicitly joins MLOps and GenAIOps into one operational discipline: Azure Machine Learning workspaces, assets, registries, training, endpoints, monitoring, Microsoft Foundry, source control, infrastructure as code, evaluation, and observability. The reference diagram can look clean while the real system depends on identity, network isolation, storage, secrets, compute, promotion paths, and ownership.
The broad capabilities in Azure machine learning services are only the starting point. A production design has to decide where experimentation ends, where reusable assets live, how environments differ, how models are promoted, which compute is trusted, and how operators know a new release is healthy.
A useful scenario is a fraud model developed by a data-science team, retrained weekly, deployed to an online endpoint, and consumed by a customer-facing application. The model is one artifact in a larger system whose security and availability depend on boundaries established long before the first deployment.
Define the workspace boundary from ownership and risk
A workspace groups jobs, compute, datastores, assets, endpoints, monitoring, and access controls. The first decision is whether one workspace can safely serve experimentation and production or whether development, test, and production need separate administrative and network boundaries.
Separate workspaces cost more operational effort and can reduce blast radius. Shared workspaces simplify reuse and can make privileged production resources visible to people who only need experimentation. The boundary should follow ownership, data sensitivity, deployment authority, and incident impact.
Workspace strategy should include lifecycle and cost as well as access. Experimentation generates notebooks, temporary compute, intermediate data, and many model candidates; production needs stable ownership, longer retention, monitored endpoints, and controlled change. Mixing those lifecycles can leave expensive idle compute or temporary assets beside critical services. Separate boundaries are valuable when they let cleanup, policy, and billing follow the real purpose of the environment instead of treating every resource as equally permanent.
Identity and secrets should be designed before pipelines
Azure RBAC and managed identities should define who can create jobs, read data, use compute, deploy models, and administer endpoints. The broader principles of Azure role-based access control matter because MLOps identities include people, training jobs, deployment identities, CI/CD workflows, and monitoring jobs.
Static credentials copied into notebooks or pipeline variables undermine that model. Centralized secrets management is safer when unavoidable secrets are stored, rotated, and accessed through narrowly scoped identities rather than embedded in source or environment files.
Machine identities deserve the same review as human roles. Training compute may need read access to curated data and write access to experiment artifacts, while deployment identities may need model, storage, and downstream-service access but no permission to create new compute. CI/CD identities may be allowed to promote tested infrastructure without reading sensitive training data. Narrowing these roles reduces the damage from a compromised pipeline and makes authorization failures easier to diagnose.
Network isolation changes what training can reach
A production workspace often needs private access to storage, registries, key vaults, and internal data. The networking concepts behind Azure virtual networks become operational because managed networks and private endpoints restrict both inbound access and what managed compute can reach outbound.
The stricter the network, the more dependencies must be enumerated: package feeds, container images, data sources, model registries, telemetry, and management APIs. Allow-only-approved outbound can reduce exfiltration risk and increases the cost of every undocumented internet dependency.
Network isolation should be planned around package and artifact supply chains. Training code often downloads Python packages, pretrained weights, container layers, or external data. A workspace moved to approved-only outbound will expose these hidden dependencies immediately. Decide whether dependencies are mirrored internally, explicitly allowed, or removed. The secure design is not one that blocks the internet abstractly; it is one where every necessary outbound relationship has an owner and a reviewable reason.
Compute and data are environment-specific even when assets are reusable
A registry can share components, environments, and models across workspaces, while jobs and endpoints remain workspace-specific resources. This separation is useful because production compute, data locations, quotas, and network rules should not be copied blindly from development.
Training reproducibility therefore needs more than a model artifact. Code, component versions, environment image, data reference, feature logic, parameters, and job metadata must identify what actually ran. Reuse is safe when environment-specific state is explicit rather than hidden inside reusable assets.
Environment-specific data should also be prevented from leaking back into reusable assets. A component promoted from development should not contain a hard-coded test storage URI or service principal. Similarly, production data references should not be copied into registry metadata that broad development users can browse. Reusability is strongest when assets define interfaces while workspaces bind them to local data, compute, and identities at execution time.
Infrastructure as code protects the workspace from drift
The current AI-300 scope includes Bicep, Azure CLI, GitHub Actions, and source control. The logic behind infrastructure as code applies directly: a workspace, compute policy, network settings, identities, and supporting resources should be reproducible enough that test and production are intentionally different rather than accidentally different.
IaC does not eliminate manual emergency changes. It gives the team a known source of intended state and a reconciliation path after incidents. If operators routinely change production settings in the portal without updating source, the architecture has two competing truths.
Infrastructure drift can be measured. Compare deployed workspace networking, role assignments, compute policy, private endpoints, diagnostic settings, and supporting resources with the source-controlled declaration. Some differences will be legitimate emergency changes; others reveal manual configuration that should be reconciled. The goal is not zero manual action during incidents but one authoritative path back to a reviewed, reproducible state after the emergency ends.
Deployment is a traffic and rollback design
Managed online endpoints can host multiple deployments and support progressive traffic movement or mirrored traffic. That makes deployment safer when teams define success metrics before shifting users and preserve the previous deployment long enough to reverse safely.
A release should connect model version, environment, scoring code, endpoint deployment, traffic percentage, approval, and monitoring. If the team knows which model is registered but not which deployment is currently serving ninety percent of requests, lifecycle traceability is incomplete.
Release traffic should be coupled with model and application compatibility. A new model can return the same schema and still shift score distribution enough to break business thresholds downstream. Before increasing traffic, test not only HTTP health and latency but the consuming application’s decision logic, fallback behavior, and any calibrated thresholds. Progressive delivery is valuable because it limits the population exposed while these integration assumptions are verified under real data.
Monitoring needs model signals and service signals
Operational telemetry such as latency, failures, CPU/memory, request volume, and endpoint availability belongs beside model-quality monitoring. The general practice in logging and monitoring on Azure helps separate infrastructure degradation from statistical drift.
A stable latency chart does not prove the model is still useful. Data drift, prediction drift, data quality, feature attribution drift, and model performance address different questions. The organization should choose signals from failure modes, not enable every metric by default.
Monitoring should also include freshness of the dependencies feeding the model. A stable endpoint can serve stale features or a delayed reference table perfectly. Expose source watermark, feature-pipeline completion, and model version alongside endpoint health so operators can distinguish scoring availability from decision freshness. User-facing applications often care more about whether the input represents today than whether the endpoint returned in 80 milliseconds.
Ownership must cross data science and operations
Data scientists understand model behavior and reference data; platform teams understand identity, compute, networking, and deployment; application owners understand latency and business integration; risk teams may define acceptable model behavior. Production operation needs all four perspectives.
Define who can register, promote, deploy, roll back, alter monitoring thresholds, approve network access, and declare a model unhealthy. Ambiguous ownership turns alerts into meetings rather than action.
Ownership boundaries should be written into escalation paths. A data scientist may own drift interpretation, a platform team may own private networking, an application team may own request latency and retries, and a risk team may own approval thresholds. Alerts should route to the team able to change the failing layer. Otherwise every incident begins with several teams proving that their component is healthy instead of one team following a known dependency chain.
A production architecture should degrade predictably
Test dependency loss: registry unavailable, training data delayed, private endpoint broken, scoring deployment unhealthy, monitoring job late, identity role removed, or secret rotated. Decide which workloads fail closed, which can use the last known-good model, and which require manual intervention.
The goal of AI operations is not to eliminate change. It is to make change traceable, reversible, observable, and owned. A clean reference diagram is useful only when the team can explain what happens to the system after one dependency in that diagram stops behaving normally.
A production readiness review should include rebuild as well as failover. Could the organization recreate the workspace, identities, network controls, registry references, and endpoint configuration in another environment if the resource group were lost or corrupted? Recovery from logical misconfiguration may require rebuilding cleanly rather than restoring one broken workspace. IaC, registries, and versioned release metadata make that path credible only when regularly tested.
Recovery testing should include a clean deployment into a second workspace or subscription boundary where policy permits. If the model, registry assets, network rules, managed identities, and IaC cannot recreate the service without the original workspace’s hidden state, the architecture remains dependent on configuration that is not actually controlled.