Amazon ECS capacity providers connect task placement with the compute supply that runs those tasks. Instead of choosing a launch type directly, a service or RunTask request can use a capacity provider strategy that distributes tasks across Fargate, Fargate Spot, or one or more Auto Scaling group capacity providers. For EC2-backed clusters, managed scaling can let ECS adjust the Auto Scaling group’s desired capacity based on pending task requirements.
Within AWS Architecture and Operations, capacity providers are the control point between application demand and infrastructure scaling. The existing ECS task placement article explains placement constraints; capacity providers determine which capacity pool gets the task and how that pool grows.
A good design separates task strategy, instance supply, interruption tolerance, and scale-in safety instead of treating “ECS autoscaling” as one feature.
Capacity provider strategy uses base first, then weight
Only one provider in a strategy can define base, which is the minimum number of tasks placed there before weighting begins.
Additional tasks are distributed according to relative weight. A 1:4 weight ratio aims for roughly one task on provider A for every four on provider B after base is satisfied.
At least one provider must have weight greater than zero.
Console and API defaults for weight differ
AWS currently documents a subtle difference: the console defaults an unspecified weight to 1, while API/CLI defaults can be 0.
Infrastructure code should always set weights explicitly rather than depending on interface defaults.
This avoids a migration where a strategy that worked when created manually becomes unusable when recreated through CLI/IaC with all weights zero.
Fargate and Fargate Spot can share one strategy
Fargate Spot uses spare capacity at a discount and can be interrupted with a two-minute warning.
A common pattern puts a base on FARGATE for guaranteed baseline service and gives additional weight to FARGATE_SPOT for interruption-tolerant burst capacity.
Applications must handle task termination and retry; Fargate does not automatically replace unavailable Spot capacity with on-demand merely because the service is under desired count.
Auto Scaling group providers let ECS drive EC2 supply
For EC2 workloads, a capacity provider references an Auto Scaling group and can enable managed scaling.
ECS groups pending tasks by resource requirements, estimates needed instances, and updates the Auto Scaling group within configured scaling step/warmup behavior.
Do not attach unrelated instances or manual scaling logic that fights the ECS-managed desired capacity.
Managed scaling target capacity controls spare headroom
Target capacity represents how full ECS intends the Auto Scaling group to be from a task-placement perspective.
Using a target below 100% keeps spare EC2 capacity available for faster task placement; 100% aims for denser utilization.
The right value depends on startup latency, burst tolerance, instance cost, and whether the workload can wait for new nodes.
Managed termination protection prevents scale-in from killing busy instances
When enabled with managed scaling and Auto Scaling instance protection, ECS protects instances that are running non-daemon tasks from Auto Scaling scale-in.
This reduces abrupt task loss during infrastructure contraction.
Use managed draining as well where appropriate so instances move through graceful task evacuation instead of disappearing under active services.
Task resource shape determines whether scale-out can succeed
A pending task may require CPU, memory, GPU, ENI, or ports that no instance type in the Auto Scaling group can provide.
Adding more undersized instances does not solve that placement problem.
Monitor provisioning tasks and instance-type compatibility; use separate capacity providers for materially different hardware classes such as GPU versus general compute.
A strategy cannot mix Fargate and ASG providers in one task request
A cluster can contain both Fargate and Auto Scaling group capacity providers, but AWS documents that one capacity provider strategy can contain only one family at a time.
Separate services or deployment variants if some workloads must run on Fargate and others on EC2.
ECS versus EKS also helps frame when this managed placement model is enough versus when Kubernetes scheduling is required.
Service migration should be deliberate
Existing ECS services can be moved from launch type to a capacity provider strategy, with deployment behavior depending on update path.
Test the new strategy in a lower environment and force/redeploy as required so existing tasks gradually land on the intended capacity pools.
Do not change both task definition resources and capacity strategy simultaneously if you want to understand which change caused placement failures.
Observe both service demand and capacity-provider supply
Monitor desired/running/pending task counts, provisioning duration, Auto Scaling group desired/in-service capacity, instance warmup, Spot interruptions, drain events, and placement failures.
A service can be “autoscaling correctly” at the task layer while capacity-provider scale-out is the actual bottleneck.
Alert on persistent provisioning state and on capacity-provider scaling that cannot satisfy pending tasks.
Capacity providers succeed when task placement and infrastructure scaling speak the same language
The mature ECS design uses explicit base/weights, distinct hardware pools, managed scaling, termination protection/draining, and interruption-aware Fargate Spot policy.
Capacity providers should turn pending task demand into predictable compute supply without hiding whether the bottleneck is task policy, instance shape, Spot availability, or Auto Scaling.
Capacity-provider design should separate heterogeneous instance families. If one Auto Scaling group contains instance types with very different CPU, memory, GPU, or ENI characteristics, ECS managed-scaling calculations and task placement can become harder to predict. Use distinct providers for hardware classes that serve different task definitions or cost/performance objectives.
Managed scaling depends on the Auto Scaling group’s configuration being compatible with ECS. AWS notes that instance weighting isn’t supported for ASGs used by ECS capacity providers. Keep launch templates and instance requirements aligned with the tasks the provider is meant to host, and let ECS manage desired count when managed scaling is enabled.
Scale-out speed is influenced by minimum/maximum scaling step sizes and instance warmup. A tiny step size can leave many tasks in PROVISIONING during bursts; an overly large step can create temporary excess EC2 capacity. Tune these values from launch time, task startup latency, and burst size rather than using one default across every cluster.
Spot-backed EC2 capacity providers need interruption handling just like Fargate Spot. Use capacity-optimized Spot strategies, multiple instance types/Availability Zones, ECS draining, and application-level retries so one reclaimed instance does not cause prolonged service degradation. A weighted strategy can reserve some on-demand baseline capacity while using Spot for elastic work.
Cluster default capacity-provider strategy is convenient but should not hide per-service intent. Services with different reliability or hardware needs should set their own strategy explicitly rather than inheriting a cluster default designed for another workload. Document which services are allowed to rely on the default and why.
Autoscaling at the ECS service layer and capacity-provider layer should be tuned together. Service Auto Scaling increases desired tasks; the capacity provider then supplies instances. If service scaling reacts in seconds but EC2 boot/warmup takes minutes, users can see latency before the infrastructure catches up. Maintain buffer capacity or slower application scaling where required.
Deployment behavior matters during strategy changes. ECS can update the capacity-provider strategy without necessarily triggering the same kind of deployment as a task-definition change, so verify where existing tasks stay and when new tasks move to the new provider. Use a controlled redeploy when you need the fleet redistributed promptly.
Cost dashboards should attribute both tasks and underlying EC2 capacity to providers. Managed scaling can leave spare headroom by design, so idle-instance cost is not always waste. Compare cost with placement latency, Spot interruption rate, and service SLO to choose an appropriate target capacity rather than optimizing utilization in isolation.
Capacity-provider strategies should be reviewed during incident failover. If a service normally uses mostly Spot capacity and Spot becomes scarce, the weights will not automatically turn Fargate Spot or EC2 Spot into on-demand capacity. Maintain enough baseline on-demand provider weight/base or an explicit operational strategy switch for workloads that cannot wait for Spot to return.
Managed draining deserves validation with deployment and termination grace periods. An instance selected for scale-in should stop receiving new tasks, drain existing tasks, and give containers enough time to handle SIGTERM and deregister from load balancers. Long-running or stateful tasks may need separate providers or protection so infrastructure scale-in does not violate service guarantees.
Task placement failures should be surfaced with the exact unsatisfied resource: CPU, memory, GPU, ENI, port, attribute, or capacity-provider availability. Pending-task alarms that only say ‘service below desired count’ force operators to rediscover whether the problem is compute supply or task definition.
Capacity-provider governance should include who can change base/weights. A small weight edit can move most of a fleet from on-demand to Spot or from one EC2 pool to another. Treat strategy changes as production routing changes with cost and resilience implications, not as harmless scheduler tuning.
Multi-provider strategies should be tested with uneven instance availability. One EC2 provider can be capacity-constrained while another remains healthy; weights express desired placement, not guaranteed capacity substitution in every scenario. Understand how pending tasks behave when their selected provider cannot launch suitable instances and whether an operational strategy update or broader instance diversification is required.
Keep provider lifecycle separate from service lifecycle. Before deleting or disassociating a capacity provider, identify every cluster default, ECS service, and RunTask workflow that references it. A provider can look idle in one dashboard while batch jobs or disaster-recovery services still depend on the name.