Build1 publisherNot yet confirmed elsewhere2 min readPublished
ECS earlySuccessCriteria ends automatic rollback once a partial rolling deployment reports success
Amazon ECS's earlySuccessCriteria can close a rolling deployment at 90 of 100 healthy tasks, ending automatic rollback for the service at that point. EventBridge and CloudTrail log that early finish exactly like a full rollout, so anything reading those events cannot tell a 90-task fleet from a 100-task one.
The Engineer · Build desk

What happened
- The ten tasks left over in that example launch afterwards through regular service scaling, outside the deployment lifecycle.
- The option covers only the ROLLING strategy, in all commercial Regions and GovCloud (US); blue/green, linear and canary deployments handle the same problem by shifting traffic.
- The healthyPercent threshold rounds up, so a desired count of 3 at 50% completes at 2 tasks and a count of 10 at 80% completes at 8.
- In DEFERRED cleanup mode ECS keeps trying to drain old-revision tasks for up to two weeks after success, and a task protected beyond that window is never cleaned up.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A replica service left at the default 100% minimumHealthyPercent cannot finish early until that setting is lowered, because healthyPercent has no other legal value there.
- exposure In the 90-of-100 setup a tenth of the fleet starts with no circuit breaker or alarm rollback behind it, so a fault that appears only in those tasks has to be caught by the operator.
- decision A CI/CD pipeline that releases its next deploy on the SUCCESSFUL event will now start that deploy while the previous revision is still filling in its last tasks.
- constraint Rounding up leaves small services with little to gain from the option; the dev.to post puts the gain there at zero.
Before September 4, 2026 [1], a rolling deployment completed only when four conditions held at once. The target revision had to reach 100% of the desired count with every task running and healthy [5]. No circuit breaker or CloudWatch alarm could have triggered a rollback, the alarm bake time had to elapse, and the source revision's tasks had to be cleaned up [5]. The dev.to post that walks through the change calls that an all-or-nothing contract, designed for a homogeneous fleet that launches in seconds [6].
The gate cannot go to zero. Before the threshold is evaluated at all, at least one target-revision task has to be launched and reported healthy [8]. Up to the gate, the circuit breaker can still roll the deployment back, and the alarm bake time still runs before completion [9]. AWS states the cost of crossing the gate directly in its documentation, according to the post [2].
No new deployment statuses were added. The lifecycle still runs IN_PROGRESS to SUCCESSFUL [12]. After success, DescribeServiceDeployments returns a snapshot with frozen counters, and DescribeServices is the only source of live counts [13]. In the 90-of-100 example, a check that wants to know whether tasks 91 to 100 ever came up has to poll DescribeServices, because the deployment record stops counting at 90 [21].
Cleanup is split the same way. BLOCKING mode drains the source revision before ECS reports SUCCESSFUL. DEFERRED reports first and drains afterward [14]. The post's second use case is services with long-lived connections, whose old revision drains slowly and holds the deployment open for reasons that have nothing to do with the new code [17]. Source-revision tasks can hold scale-in protection for up to 2,880 minutes [16], or 48 hours [22].
The post's motivating case is GPU-accelerated inference. The tail of a fleet on specialized capacity launches at the pace of hardware availability, not the container, so the deployment stays open for two tasks and the deploys queued behind it never start [18]. For that service I would turn the option on. The pairing I'd use is BLOCKING cleanup, so that SUCCESSFUL at least means the old revision is gone [14]. I'd also add a DescribeServices check that alerts when running tasks stay below the desired count after success [13]. For a service with long-lived connections I'd look at DEFERRED instead and track the old revision's task count down to zero myself [14].
What to watch
- Whether AWS adds a distinct status or event field that marks an early success in EventBridge or CloudTrail.
- Whether earlySuccessCriteria is extended beyond ROLLING to the BLUE_GREEN, LINEAR or CANARY strategies.
- Whether DescribeServiceDeployments starts reporting live counts after success, so the deployment record shows the late tasks arriving.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Since September 4, 2026, Amazon ECS earlySuccessCriteria lets you close a rolling deployment once a fraction of tasks is healthy.
- [2]
The ECS documentation says that after the deployment completes, the circuit breaker and alarm-based rollback no longer roll the service back.
- [3]
With healthyPercent 90 and a desired count of 100, the deployment turns SUCCESSFUL at task 90; the remaining ten launch afterwards through regular service scaling, outside the deployment lifecycle.
- [4]
earlySuccessCriteria is available for the ROLLING strategy in all commercial Regions and GovCloud (US); the BLUE_GREEN, LINEAR and CANARY strategies have no earlySuccessCriteria and solve the same problem through traffic shifting, at a different cost.
- [5]
Until now, an ECS rolling deployment only completed when four things were true at once: the target revision reached 100% of the desired count with every task running and healthy, neither the circuit breaker nor a CloudWatch alarm triggered a rollback, the alarm bake time elapsed, and the source revision tasks were cleaned up.
- [6]
The dev.to post describes the old completion rule as an all-or-nothing contract designed for a homogeneous fleet that launches in seconds.
- [7]
healthyPercent is an integer between 0 and 100 and must sit between the service minimumHealthyPercent and 100; with the 100% default for replica services, there is no range left.
- [8]
ECS launches at least one task on the target revision and waits for it to be healthy; the first healthy task is a precondition of the evaluation.
- [9]
The deployment circuit breaker can roll back up to the success gate, and the CloudWatch alarm bake time elapses before completion.
- [10]
healthyPercent rounds up: a desired count of 3 at 50% completes at 2 tasks, not 1, and a desired count of 10 at 80% completes at 8.
- [11]
According to the dev.to post, on a small service the gain from earlySuccessCriteria is zero.
- [12]
earlySuccessCriteria adds no new deployment statuses; the lifecycle stays IN_PROGRESS to SUCCESSFUL, and nothing in EventBridge or CloudTrail tells an early success apart from a full one.
- [13]
After success, DescribeServiceDeployments returns a snapshot that freezes the counters; DescribeServices, with live counts, is the only source of real counts.
- [14]
In BLOCKING cleanup mode ECS drains source revision tasks before reporting SUCCESSFUL; in DEFERRED mode it reports SUCCESSFUL before draining.
- [15]
In DEFERRED mode, ECS tries to drain source revision tasks for up to two weeks after success; a task protected beyond that is never cleaned up.
- [16]
Source revision tasks can carry scale-in protection of up to 2880 minutes, according to the post's diagram.
- [17]
Services with long-lived connections have a source revision that is slow to drain and holds the deployment open for a reason that has nothing to do with the health of the new code.
- [18]
The motivating case is GPU-accelerated inference: the tail of a fleet on specialized capacity does not launch at the speed of its head, the variable is hardware availability rather than the container, and the deployment stays open waiting for two tasks while the pipeline queues and the next deploy never starts.
- [19]
The SUCCESSFUL event in EventBridge unblocks the next deploy in the CI/CD pipeline, according to the post's diagram.
- [20]
On a replica service at the default minimumHealthyPercent of 100, the only valid healthyPercent is 100, so early success cannot happen until minimumHealthyPercent is lowered.
- [21]
In the 90-of-100 example, whether the remaining ten tasks came up is visible only through DescribeServices, because DescribeServiceDeployments freezes its counters at success.
- [22]
The 2,880-minute scale-in protection cap equals 48 hours.
- [23]
In the 90-of-100 example, 10 of 100 tasks, or 10% of the desired count, launch after automatic rollback has stopped covering the service.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toECS Early Success Criteria: where your deployment rollback ends
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Deployment rollbackFollow
- CI/CD pipelinesFollow
- Container OrchestrationFollow