Azure compute decisions look simple when they are reduced to a list: virtual machines, app platforms, containers, functions, and other managed runtimes. The difficult part is not knowing that these options exist. It is deciding how much operating control the workload needs, how quickly it must scale, how state is handled, what deployment model the team can support, and which constraints make an apparently modern option a poor fit. Those judgment calls sit behind the compute section of the current AZ-900 foundation.
A useful comparison begins with responsibility. A virtual machine gives the team broad control over the operating system and installed software, but that control also brings patching, image management, configuration, capacity, and security work. A managed application platform removes some of that burden but introduces platform conventions. Serverless execution can reduce idle capacity but depends on event-driven design and service limits. Containers create portability and packaging benefits while still requiring an orchestration and operations model.
The Azure Fundamentals credential should make this reasoning easier rather than encourage a feature-by-feature winner. There is no universal best compute service. The best option is the one whose constraints and responsibilities match the workload and the team that must operate it.
A concrete example helps expose hidden criteria. Suppose a company has a .NET line-of-business application that runs continuously, stores uploaded files locally, uses Windows-integrated dependencies, and receives predictable daytime traffic. Moving it directly to functions would require a major redesign of state and execution assumptions. App Service might reduce operating-system work but still require changes around file persistence and identity. Virtual machines could provide the fastest low-risk migration while the team gradually externalizes state. The “least cloud-native” first step can be the most rational when migration risk and reversibility matter.
Now compare a new image-processing workflow triggered whenever a file arrives. It has no long-lived session, can retry safely, and runs for seconds or minutes. Keeping dedicated VMs online for that pattern would create idle cost and unnecessary operational work. Event-driven or container-based compute may fit because the workload is naturally decomposable and execution can scale with arrivals. The same organization can therefore make different compute choices without inconsistency.
Operations maturity should influence the decision as well. Kubernetes can provide powerful scheduling and portability, but adopting it for a small application can create a control plane, deployment, observability, networking, and skills burden the team did not previously have. A managed application platform may offer less low-level control and far less day-two complexity. Architecture should optimize the whole operating system, not the prestige of the runtime.
Finally, define an exit condition. If a managed service becomes too expensive at sustained scale, what would moving look like? If a VM-hosted application needs faster release velocity, what prerequisites must be removed before migration? Recording those triggers makes the original choice easier to defend because the team is not pretending it will remain correct forever.
This decision becomes easier when teams score the constraints in writing. A one-page record listing runtime requirements, state, traffic shape, networking, deployment model, security ownership, and cost sensitivity is often more useful than a long service comparison. It preserves why the choice was made and gives future reviewers a clear trigger for reconsidering it.
Start with the shape of the workload
Before discussing services, describe the workload. Is it a continuously running web application, a scheduled batch job, an event-driven integration, a stateful legacy service, a short-lived API, or a data-processing task? Does demand stay stable, follow a daily curve, or arrive in unpredictable bursts? The workload shape determines which scaling and billing characteristics matter.
A steady workload can make reserved or continuously running capacity reasonable. A bursty workload may benefit from automatic scaling or consumption-based execution. A long-running stateful process may be awkward on an execution model designed for short events. Matching workload behavior to compute behavior is more important than choosing the newest abstraction.
Decide how much operating-system control is genuinely required
Teams often default to virtual machines because they are familiar and flexible. The key question is whether the application actually requires operating-system control. Custom drivers, legacy dependencies, specialized agents, unusual network behavior, or unsupported runtimes may justify it. If those constraints do not exist, carrying an operating system can add maintenance without adding business value.
Control also creates accountability. If the team owns the guest operating system, it owns patch timing, hardening, malware protection, configuration drift, and recovery of that layer. The decision should therefore include who will perform those tasks, how reliably they can do so, and whether a managed service can remove work the organization does not need to differentiate on.
Scaling model should match demand, not ambition
“It needs to scale” is too vague to guide compute selection. Define whether the workload needs vertical scaling, horizontal scaling, rapid burst capacity, scheduled expansion, or simply predictable headroom. Some services make horizontal scaling straightforward when applications are stateless. Others are better for long-running processes or workloads whose state cannot be easily separated from the instance.
The architecture should also define how quickly scale must occur and what happens while it catches up. A workload that can queue requests has different needs from one that must answer synchronously within milliseconds. Scaling is a behavior of the full system, including data stores, downstream APIs, and network paths. Adding compute capacity is useless if another dependency remains the bottleneck.
State changes the answer
State is one of the most important compute constraints. Local session data, files on the instance, in-memory coordination, and tightly coupled databases can make horizontal scaling or instance replacement difficult. Managed compute options work best when state is placed in services designed to persist independently of the execution instance.
A legacy application may need to remain stateful for valid reasons. In that case, forcing it into a stateless pattern can create migration risk without enough benefit. The better decision may be to stabilize it on virtual machines first, externalize selected state over time, and only then move to a more managed runtime. Reversibility and migration sequence belong in the compute decision.
Deployment frequency affects platform fit
A team that deploys many times per day values repeatable builds, rapid rollback, isolated releases, and automated deployment. A rarely changed line-of-business application may value stability and compatibility more than deployment velocity. Compute platforms differ in how naturally they support immutable images, deployment slots, revisions, or infrastructure-defined environments.
The Azure compute development patterns are useful because they show how application architecture and deployment model influence each other. Choose a platform that supports the team’s real release process, not an imagined future process nobody has yet built.
Networking requirements can narrow the options
Private connectivity, inbound exposure, outbound control, hybrid dependencies, service endpoints, network appliances, and custom protocols can all affect compute choice. A managed platform may simplify application hosting while requiring a specific integration approach for virtual networks or private services. A VM may provide broad network flexibility but also increase the configuration surface.
Map required flows before selecting the runtime. Identify which connections are inbound, outbound, east-west, administrative, or data-plane. Determine whether the workload must reach on-premises systems, whether public endpoints are acceptable, and how DNS and certificates are managed. Compute cannot be evaluated independently of connectivity.
Failure behavior matters more than restart behavior
Every compute service can fail. The architecture question is what a failure does to user work and how the service recovers. If an instance disappears, is another instance ready? Is state preserved elsewhere? Does traffic stop using the unhealthy endpoint? Can a deployment failure be rolled back without data loss? Those questions distinguish real resilience from a service that merely restarts.
Test partial failure rather than only total outage. One instance can be unhealthy while the platform considers it running. A downstream dependency can time out and cause thread or connection exhaustion. A region can remain online while a specific service is degraded. Compute architecture should define health signals that represent the application outcome, not just infrastructure status.
Cost should include idle capacity and operating effort
Direct price comparisons are often misleading because services bill differently and shift labor differently. Virtual machines may look inexpensive per hour but carry management work. Serverless execution may be efficient at low or intermittent volume but expensive under sustained patterns. Managed platforms can cost more per unit while reducing operational burden and deployment effort.
Measure the expected workload over time. Include idle periods, peak periods, nonproduction environments, scaling buffers, logging, backups, and network traffic. Then include the human work required to maintain the platform. The cheapest compute line item is not necessarily the least expensive system.
Use reversibility as a design criterion
Compute choices create different levels of coupling. A standard virtual machine image can be portable but leaves more work with the team. A deeply integrated managed platform can reduce operations while making migration require application changes. Neither outcome is inherently bad. The team should know which parts of the design are easy to reverse and which create long-lived commitments.
The broader Microsoft Azure ecosystem offers many compute abstractions because workloads differ. A mature decision can explain the workload shape, control requirements, state model, scaling behavior, connectivity, operations, cost, and exit path. Once those criteria are explicit, the right compute option usually becomes much easier to defend.