BuildNot yet confirmed elsewhere1 publisher3 min readPublished
A 28-host Debian 12 cutover, and the 02:13 failure Ansible could not have prevented
A team's Debian 11-to-12 move across 28 hosts put the usual Ansible complaints under load. The failure that actually stopped the run sat in the SSH layer the tool does not own.
The Engineer · Build desk
What happened
- Twenty-eight application and utility hosts went to Debian 12 over three evenings, among them 14 API nodes and six Celery workers.
- The rerun after the fix reported 23 hosts ok, four with expected package changes, and one failing on a stale internal repository key.
- The author rates the "just SSH" complaint false and the belief that idempotence makes reruns safe dangerously false.
Why it matters
- constraint The guarantee is scoped to the task, not to the post-install scripts and handlers around it, so a rerun needs a blast-radius review rather than a green summary line.
- exposure The layer that stranded a payments host mid-run is the one config management does not administer, which means a cipher or key policy change can lock the tool out of exactly the hosts it was booked...
- decision Trading Restart=always for Restart=on-failure buys quiet at the price of self-healing, so someone has to own an alert for a worker that simply stays down.
Nothing in the play caused the stop at 02:13. A Debian 11 host in the payments group refused the old OpenSSH cipher list [5], which is a property of sshd and the client config on that pair of machines, not of any task in the run. Work backwards from the 41 quiet minutes and the run began around 01:32 [16]. The write-up never says what moved first, and that gap is the interesting part: during a distribution cutover, the transport you automate over is one of the things being upgraded.
The rerun tally is the number worth keeping. Twenty-three hosts ok, four changed, one failed [1] comes to 28 [14], which is the whole fleet [4]. The second pass therefore reached every host including the one that had dropped out, and the only failure left was a stale internal repository key [1]: content, in the play, fixable by whoever owns the key. The unrecoverable class of failure was the transport, and the transport is the part ansible-core 2.17.7 in a pinned virtualenv [6] has no authority over.
The month before the cutover says the same thing. Four separate ways to lose a connection before a single task runs: a stale known_hosts entry, a rotated host key, an expired jump-host certificate, an orphaned control socket under /tmp [2]. Call it one new class of transport failure per week [15], with no distro upgrade to blame. That is the sense in which "it's just SSH" is wrong and still worth respecting. Inventory groups, facts, handlers, check mode and a run history you can query afterwards are all real [8], and the APT source went in through deb822_repository with nginx reloaded only when the file changed [13], which no shell loop gives you across hosts that differ as much as gunicorn nodes, workers with their own systemd limits, and reporting boxes holding an NFS mount [9]. None of it starts without a working socket [7].
The Celery incident belongs to a different family. Six workers restarted on a run that should have reported no changes, because a task was written with state: restarted [12]. The module was idempotent; the play author was not. That is a code review problem rather than a tool problem, and it is precisely what the "safe to rerun" belief conceals [10]. The unit change underneath it, from Restart=always to Restart=on-failure after a wrong RabbitMQ credential produced 9 GB of logs in one on-call shift [11], moves the failure mode from loud to quiet: the next bad credential leaves stopped workers instead of a log flood, and queue depth has to do the detecting.
One more thing about the source itself. The published example breaks off at "- name: Install worker" without the corrected task [3]. On a run whose expensive lesson was a transport failure and whose self-inflicted one was a shortcut inside a handler, that missing snippet is the part everyone else would have copied.
What to watch
- Whether a follow-up publishes the corrected worker task, since the handler pattern that replaced state: restarted is the part currently missing.
- Whether the stale internal repository key was a rotation or an expiry, which decides how much of the rest of the fleet is on borrowed time.
- Whether Restart=on-failure holds up the next time a credential is wrong, or just produces a silent version of the same incident.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence34
- Adoption21
- Hype gap+7
- Incentives27
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
After fixing the cipher issue, the rerun showed 23 hosts reporting ok, four making the expected package changes, and one failing on a stale internal repository key.
- [2]
In the preceding month the team saw all four categories of SSH-side failure that can stop a run before a task begins: a stale known_hosts entry, a host key rotation, an expired jump-host certificate, and a control socket left under /tmp.
- [3]
The published write-up breaks off at the line "- name: Install worker" without showing the replacement task that removed state: restarted.
- [4]
A team moved 28 application and utility hosts to Debian 12 over three evenings: 14 API nodes, six Celery workers, four reporting boxes, two Redis replicas, and two machines nobody could name without checking NetBox.
- [5]
At 02:13, the last Debian 11 VM in the payments group stopped accepting the old OpenSSH cipher list, halfway through an Ansible run that had looked boring for 41 minutes.
- [6]
The team ran ansible-core 2.17.7 from a pinned Python 3.12 virtual environment on the control host.
- [7]
The author's verdict on the claim that Ansible is just SSH is false, though SSH is still where many bad evenings begin.
- [8]
The cutover used inventory groups, facts, idempotent modules, handlers, privilege escalation rules, check mode, variable precedence, and a run history the team could inspect afterwards.
- [9]
The hosts were not alike: API nodes run gunicorn behind Nginx, Celery workers have different systemd limits, and the reporting boxes mount an NFS share the application hosts never see.
- [10]
The author calls the belief that idempotence makes a rerun safe dangerously false: a package install can be idempotent while its post-install script restarts a service, and a template can be idempotent while its handler drains a connection pool at a bad moment.
- [11]
The Celery worker unit used Restart=always on Debian 11 and was changed to Restart=on-failure on Debian 12 after a bad RabbitMQ credential caused a restart loop and 9 GB of logs over an on-call shift.
- [12]
The second run should have reported no changes but restarted all six Celery workers because a task used state: restarted as a substitute for thinking.
- [13]
The play configured the internal APT source with ansible.builtin.deb822_repository and installed packages with ansible.builtin.apt, notifying a handler that reloaded nginx only when the repository file changed.
- [14]
The rerun outcomes account for the entire fleet: 23 ok plus four changed plus one failed equals 28 hosts.
- [15]
Four distinct classes of SSH transport failure in the month before the cutover works out to about one new class per week.
- [16]
The failing run had started at roughly 01:32, since the break came at 02:13 after 41 minutes of uneventful execution.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toAnsible Myths We Met During The Debian 12 Cutover
1 article · August 21, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.