Build1 publisher3 min readPublished
Two half-opens are not one state: the ABA bug that gated breakwater 1.0
A shared circuit breaker looks like a three-state machine until a stale decision lands on a fresh half-open period and kills a recovery it never observed.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- breakwater is a resilience toolkit for Node.js covering retry, circuit breaker, timeout, bulkhead, rate limiting and stale-while-open caching, composable, with observability built in.
- breakwater has released version 1.0.0, and its headline feature is a circuit breaker whose state is shared across every instance of a service.
- Intended behaviour of the shared breaker: one instance sees the outage and trips the breaker, other instances fail fast immediately without each discovering the outage on their own users, and when the cooldown elapses exactly one instance probes the recovering dependency while the rest keep waiting.
- A circuit breaker is described as a tiny state machine with three states: closed, open and half-open.
- In a single process the author protects state transitions with nothing at all, because JavaScript is single-threaded and the code between two awaits cannot be interrupted.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
breakwater, a Node.js resilience toolkit covering retry, circuit breaker, timeout, bulkhead, rate limiting and stale-while-open caching, has shipped 1.0.0, and its headline feature is a circuit breaker whose state is shared across every instance of a service [1][2]. According to its author, the release was gated by a correctness bug with a general lesson: once state is shared, closed/open/half-open is not a sufficient state space, because half-open entered twice is not the same state [4][8].
The intended behaviour is straightforward. One instance sees the outage and trips the breaker, the others fail fast without rediscovering the outage on their own users, and when the cooldown elapses exactly one instance probes the recovering dependency while the rest keep waiting [3].
In a single process, none of the transitions need protection: JavaScript is single-threaded, and the code between two awaits cannot be interrupted [5]. Shared across N instances, every transition becomes a compare-and-set, which the first store interface expressed as `transition(name, from, to): boolean`, swapping only if the state was still `from`, implemented atomically in Redis with a Lua script [6].
The author's failing sequence: a probe fails and the instance decides to reopen; that decision spends a few milliseconds in a Lua round trip; meanwhile another probe succeeds, reaches the majority, closes the circuit, traffic resumes, fails again, reopens, waits out the cooldown and re-enters half-open. The first instance's swap then lands, asks whether the state is still half-open, finds that it is, and succeeds, killing a recovery it never observed on the basis of a decision belonging to a period that ended three transitions earlier [7]. This is ABA, disguised by the fact that the states have names; half-open is a label the circuit wears repeatedly, not an identity [8].
The first attempt at a fix is the part worth studying. The race predated the distributed store, since the in-process breaker had the same window whenever a custom store was async [9], so the author patched it: after a successful swap, check whether the period had flipped and, if so, swap the state back, with a comment describing the behaviour as best effort until stores could fence the CAS with a generation [10]. His own verdict is that two swaps are not one swap, another instance sees the wrong state in between, the compensating swap can itself fail, and he shipped it only because the alternative was redesigning the store contract [11].
The eventual fix makes the generation explicit: a monotonic fence token minted on every successful transition, `readState` returning state, fence and optional openedAt, and `compareAndSet(name, from, to, fence)` returning ok plus a snapshot [12]. A swap lands only if the state is still `from` and nothing has transitioned since the fence was read, so a stale decision carrying fence 7 against a store holding fence 10 is refused atomically inside the same script [13] - three transitions of drift, visible as arithmetic rather than as a label comparison [1]. Both compensating transitions were deleted [14].
Two consequences fell out. A lost race now returns the current snapshot, where it previously cost a second round trip to learn what had happened [15]. And openedAt now lives in Redis, stamped from the server clock, instead of each instance counting the cooldown from when it first noticed the trip, which had left instances disagreeing about when probing was allowed and let the earliest noticer probe too soon [16].
What to watch is the cost of the premise: Redis is now on the path of every protected call [17]. The author's answer is that no store method ever rejects, and that an unreachable Redis falls back to what the instance already has [17]; the supplied account breaks off there, so the thing to verify in the shipped code is whether the exactly-one-prober guarantee [3] survives a fallback that answers from local state.