Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

A 28-host Debian 12 cutover, and the 02:13 failure Ansible could not have prevented

A team's Debian 11-to-12 move across 28 hosts put the usual Ansible complaints under load. The failure that actually stopped the run sat in the SSH layer the tool does not own.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Twenty-eight application and utility hosts went to Debian 12 over three evenings, among them 14 API nodes and six Celery workers.
  • The rerun after the fix reported 23 hosts ok, four with expected package changes, and one failing on a stale internal repository key.
  • The author rates the "just SSH" complaint false and the belief that idempotence makes reruns safe dangerously false.

Why it matters

  • constraint The guarantee is scoped to the task, not to the post-install scripts and handlers around it, so a rerun needs a blast-radius review rather than a green summary line.
  • exposure The layer that stranded a payments host mid-run is the one config management does not administer, which means a cipher or key policy change can lock the tool out of exactly the hosts it was booked...
  • decision Trading Restart=always for Restart=on-failure buys quiet at the price of self-healing, so someone has to own an alert for a worker that simply stays down.

Nothing in the play caused the stop at 02:13. A Debian 11 host in the payments group refused the old OpenSSH cipher list [5], which is a property of sshd and the client config on that pair of machines, not of any task in the run. Work backwards from the 41 quiet minutes and the run began around 01:32 [16]. The write-up never says what moved first, and that gap is the interesting part: during a distribution cutover, the transport you automate over is one of the things being upgraded.

The rerun tally is the number worth keeping. Twenty-three hosts ok, four changed, one failed [1] comes to 28 [14], which is the whole fleet [4]. The second pass therefore reached every host including the one that had dropped out, and the only failure left was a stale internal repository key [1]: content, in the play, fixable by whoever owns the key. The unrecoverable class of failure was the transport, and the transport is the part ansible-core 2.17.7 in a pinned virtualenv [6] has no authority over.

The month before the cutover says the same thing. Four separate ways to lose a connection before a single task runs: a stale known_hosts entry, a rotated host key, an expired jump-host certificate, an orphaned control socket under /tmp [2]. Call it one new class of transport failure per week [15], with no distro upgrade to blame. That is the sense in which "it's just SSH" is wrong and still worth respecting. Inventory groups, facts, handlers, check mode and a run history you can query afterwards are all real [8], and the APT source went in through deb822_repository with nginx reloaded only when the file changed [13], which no shell loop gives you across hosts that differ as much as gunicorn nodes, workers with their own systemd limits, and reporting boxes holding an NFS mount [9]. None of it starts without a working socket [7].

The Celery incident belongs to a different family. Six workers restarted on a run that should have reported no changes, because a task was written with state: restarted [12]. The module was idempotent; the play author was not. That is a code review problem rather than a tool problem, and it is precisely what the "safe to rerun" belief conceals [10]. The unit change underneath it, from Restart=always to Restart=on-failure after a wrong RabbitMQ credential produced 9 GB of logs in one on-call shift [11], moves the failure mode from loud to quiet: the next bad credential leaves stopped workers instead of a log flood, and queue depth has to do the detecting.

One more thing about the source itself. The published example breaks off at "- name: Install worker" without the corrected task [3]. On a run whose expensive lesson was a transport failure and whose self-inflicted one was a shortcut inside a handler, that missing snippet is the part everyone else would have copied.

What to watch

  • Whether a follow-up publishes the corrected worker task, since the handler pattern that replaced state: restarted is the part currently missing.
  • Whether the stale internal repository key was a rotation or an expiry, which decides how much of the rest of the fleet is on borrowed time.
  • Whether Restart=on-failure holds up the next time a credential is wrong, or just produces a silent version of the same incident.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence34
Adoption21
Hype gap+7
Incentives27
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    After fixing the cipher issue, the rerun showed 23 hosts reporting ok, four making the expected package changes, and one failing on a stale internal repository key.

  2. [2]

    In the preceding month the team saw all four categories of SSH-side failure that can stop a run before a task begins: a stale known_hosts entry, a host key rotation, an expired jump-host certificate, and a control socket left under /tmp.

  3. [3]

    The published write-up breaks off at the line "- name: Install worker" without showing the replacement task that removed state: restarted.

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 21, 2026

    Ansible Myths We Met During The Debian 12 Cutover

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • systemd Service Lifecycle ManagementFollow
  • Debian Major Version UpgradesFollow
  • Ansible Configuration ManagementFollow
  • SSH Transport ReliabilityFollow
  • On-Call Incident LessonsFollow
Loading related stories