Build1 publisherNot yet confirmed elsewhere3 min readPublished
Autoscaling Is the One Config Surface With Two Writers and No Merge
Scaling policies on ECS, ASG, VMSS and MIG are separate API objects with last-writer-wins semantics. That is why inherited scaling behavior stays wrong, and why nobody gets told when it changes.
The Engineer · Build desk
What happened
- Across ECS, ASG, VMSS and MIG the pattern is identical: the autoscaling policy is a separate object with multiple possible writers and last-writer-wins semantics. Nothing merges. Nothing warns.
- Teams inherit autoscaling config they did not write: the author left, the Terraform was half-applied, someone tuned a threshold in the console during an incident two years ago, and production scaling behavior is now an unknowable merge of code, clicks and defaults that everyone is afraid to touch.
- AWS ASG scaling policies and the Application Auto Scaling targets and policies used by ECS live as their own API objects; if Terraform or CloudFormation defines them, the next apply silently reverts every console tweak made since, including a load-bearing incident-era threshold change.
- Policies created by hand in the AWS console are invisible to the code, survive until someone cleans up drift, and then vanish.
- Either direction of overwrite changes the system's real behavior with no deploy, no PR and no alert.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A writeup on dev.to about adopting autoscaling config you did not write names the mechanism more precisely than most drift discussions manage: across ECS, ASG, VMSS and MIG, the scaling policy is its own object, with multiple possible writers and last-writer-wins semantics [4]. Nothing merges and nothing warns, which means the real scaling behavior of production can change with no deploy, no PR and no alert [4][7].
The specifics differ by platform only in blast radius. On AWS, ASG scaling policies and the Application Auto Scaling targets and policies behind ECS services are separate API objects; if Terraform or CloudFormation declares them, the next apply silently reverts every console tweak made since, including the incident-era threshold that turned out to be load-bearing [5]. The inverse is equally common: a policy created by hand in the console is invisible to the code, survives right up until someone cleans up drift, and then disappears [6]. Azure VMSS autoscale is a single autoscaleSettings resource containing profiles and rules, so portal edits modify it in place and the next ARM, Bicep or Terraform deployment that also defines it replaces the whole object; you do not lose one rule, you lose the entire tuned profile set at once [8]. On GCP, the MIG autoscaler is likewise its own attached object, gcloud edits and Terraform definitions overwrite each other whole, and a MIG that gets deleted and recreated comes back with whatever the code says rather than what the console said [9].
The ordering advice in the piece is the part worth stealing. Inventory from the API rather than from the repo, using describe-policies and describe-scaling-policies on AWS, az monitor autoscale list, and gcloud compute instance-groups managed list [10]. Then freeze before fixing: pick one source of truth, almost always the IaC, import the live policy objects verbatim including the warts, and diff until plan shows zero changes, so the first apply after adoption is a no-op [2]. Tuning while two writers still exist is the mistake that produces the next inherited mess.
Only once there is one honest copy do the smells become checkable, and the source lists six that inherited policies reliably contain: cooldowns under 120 seconds, which make the group oscillate on its own noise [16]; targets above 90 percent, where new capacity arrives after the damage and the policy functions as a post-incident notification system [17]; targets below 30 percent, which is permanent over-provisioning in an autoscaling costume [18]; min equal to max, which is a fixed fleet with extra steps [11]; no policy at all on a group everyone assumed had one [12]; and step policies stacked on target tracking against the same metric, two controllers on one wheel [13]. The usable target band implied by those two thresholds is 60 percentage points wide, which is a lot of room to be wrong inside [15].
Enforcement is the cheap half. CloudTrail records PutScalingPolicy and PutAutoScalingPolicy, Azure activity logs and GCP audit logs record their equivalents, so an alert on policy-write events whose principal is not the deploy pipeline converts the next silent overwrite into a loud one [3]. Incident-time tuning still works; it just arrives with a follow-up task to codify or revert instead of becoming un-owned state [14].
Watch two signals. The first plan after import should show zero changes, and if it does not, the import is not finished [2]. After that, watch for policy writes attributed to a human principal [3], and treat any tuning that changes target, cooldown and max in one apply as a lost experiment rather than a fix [19].
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+14
- Incentives22
- Confidence46
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Teams inherit autoscaling config they did not write: the author left, the Terraform was half-applied, someone tuned a threshold in the console during an incident two years ago, and production scaling behavior is now an unknowable merge of code, clicks and defaults that everyone is afraid to touch.
- [2]
Freeze before fixing: the worst adoption mistake is tuning while two writers still exist. Decide the single source of truth (almost always the IaC), export the live state into it verbatim including warts, import the live policy objects with terraform import or the ARM/gcloud equivalent, and diff until plan shows zero changes so the first apply after adoption is a no-op.
- [3]
CloudTrail records PutScalingPolicy and PutAutoScalingPolicy, and Azure activity logs and GCP audit logs record policy writes; an alert on policy-write events that did not come from the deploy pipeline's principal turns the next silent overwrite into a loud one.
- [4]
Across ECS, ASG, VMSS and MIG the pattern is identical: the autoscaling policy is a separate object with multiple possible writers and last-writer-wins semantics. Nothing merges. Nothing warns.
- [5]
AWS ASG scaling policies and the Application Auto Scaling targets and policies used by ECS live as their own API objects; if Terraform or CloudFormation defines them, the next apply silently reverts every console tweak made since, including a load-bearing incident-era threshold change.
- [6]
Policies created by hand in the AWS console are invisible to the code, survive until someone cleans up drift, and then vanish.
- [7]
Either direction of overwrite changes the system's real behavior with no deploy, no PR and no alert.
- [8]
Azure VMSS autoscale is a separate autoscaleSettings resource attached to the scale set, containing profiles and rules; portal edits modify it in place, and the next ARM, Bicep or Terraform deployment that also defines it replaces the whole object, so you lose the entire tuned profile set rather than one rule.
- [9]
On GCP the autoscaler is its own object attached to the managed instance group; gcloud edits and Terraform definitions overwrite each other whole, and a deleted-then-recreated MIG comes back with whatever the code says, not what the console said.
- [10]
Step one is to inventory every policy that exists according to the API rather than the code, using aws autoscaling describe-policies, aws application-autoscaling describe-scaling-policies --service-namespace ecs, az monitor autoscale list, and gcloud compute instance-groups managed list.
- [11]
Smell: min equal to max is not autoscaling at all, but a fixed fleet with extra steps, worth making explicit.
- [12]
Smell: no policy on a group that clearly expected one, the scaling everyone assumes exists and does not.
- [13]
Smell: step policies stacked with target tracking on the same metric, two controllers steering one wheel.
- [14]
Incident-time tuning stays possible under this model; it arrives with a follow-up task to codify or revert instead of becoming un-owned state.
- [15]
The band between the two named target thresholds, 30 percent and 90 percent, is 60 percentage points wide.
- [16]
Smell: cooldowns under 120 seconds make the group react to its own noise, oscillating up and down, which costs money on the way up and availability on the way down.
- [17]
Smell: targets above 90 percent trigger scaling so late that new capacity arrives after the damage, making the policy effectively a post-incident notification system.
- [18]
Smell: targets below 30 percent are perpetual over-provisioning wearing an autoscaling costume.
- [19]
Tune last, one variable at a time, through the pipeline, watched across a full traffic cycle; changing target, cooldown and max in one apply teaches nothing except fear if behavior degrades.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toAdopting an Autoscaling Policy You Didn't Write: ECS, ASG, VMSS and MIG Without Silent Overwrites
1 article · August 21, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.