Skip to content

Build1 publisher3 min readPublished

Squirix silenced seven analyzer warnings that showed its failover was never wired in

Squirix's maintainer says the failover promised in 0.1.0-preview.8 never runs in the shipped server, despite unit tests, E2E tests and green CI. The build's dead-code analyzer had reported the gap for over two weeks, with every warning suppressed.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Squirix silenced seven analyzer warnings that showed its failover was never wired in
Generated illustration

What happened

  • The preview.8 release notes said RF>=3 clusters survive losing a node: the majority elects a new leader, stale leaders are fenced and clients are rerouted.
  • The replication gRPC contract had no PreVote or RequestVote RPC, so nodes could not ask each other for a vote over the network.
  • Two separate tests in the release suite were meant to prove failover, and both stopped a follower instead of the key's leader.
  • Every failover issue in the milestone was closed, but the exit-gate checklist for enabling failover in production had no boxes ticked when the release was tagged.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Anyone running a preview.8 RF=3 cluster on the strength of the release notes has preview.7 availability: when a key's owner goes down, that key stays unavailable until the owner returns.
  • constraint An assertion on the value read back cannot tell a leader election from the loss of a follower, so a failover test proves nothing unless it stops the key's owner.
  • decision Teams that silence dead-code warnings with a promise of later wiring now have a documented case for tying each suppression to a tracked item the build can check.

Meziantou.Analyzer's rule MA0182 flags internal types that nothing references, and Squirix builds with it [3]. Each suppression in the replication code carried nearly the same justification [3]: "Test-only activation seam until failover activation wires the gate in a follow-up milestone." [4] The suppressed types were the election timer, the failover activation gate, the leader read barrier, the snapshot catch-up session, the reroute budget, the stale-term classifier and the election commit rule [5]. "That is most of failover," the maintainer wrote [20].

The milestone in that string did not exist. The milestone folder held preview.8 itself and a backlog with one unrelated item, and git log -S placed the suppressions inside the preview.8 pull requests around September 6 and 7 [6]. The maintainer found the gap on September 23 [2], about 16 to 17 days later [1]. "The follow-up only ever existed in the justification string," the maintainer wrote [19]. The maintainer added that the analyzer had been right the whole time: it said nobody used these types, and nobody did [7].

Every route into failover ran through test code. The AutomaticFailoverEnabled and QuorumReadsEnabled switches lived on the internal topology options. Nothing in src/ read them, and only tests set them to true [8]. The repair service was registered and running, and nothing ever queued work for it [10]. The request path had no reroute or stale-term handling, although the release notes promised "bounded client rerouting on stale terms" [11].

The milestone's hard requirement was that stopping the leader of an RF=3 group leaves the remaining majority serving reads and writes within five seconds [14]. The test for it picked a key owned by nodeB, then stopped nodeA [15]. For that key, nodeA was only a follower. The owner stayed up, nodeB and nodeC still made a majority, the write committed and the read returned it [15]. No election was needed, and none happened [15]. The majority-commit write path is one of the pieces that did ship [13], so the test exercised working code and passed well inside five seconds [15]. "This part bothers me most, because it existed to catch exactly this," the maintainer wrote [17].

The assertion compared a value. A value comes back whenever the owner survives, election or no election. In my view a failover proof has to choose its victim from the key's current owner at run time and then assert that a new leader took over. I'd also treat a suppression whose justification names future work as an open item, and fail the build when that work has no entry in the milestone plan. Going by the maintainer's account, both checks would have failed on preview.8 [6][15].

The write-up itself is careful work. It prints the test as it stood at the preview.8 tag [15] and lists what did run, including the replication opt-in with its persistence and mTLS checks, durable follower logs and the majority-commit write path [13]. The maintainer has added an update to the first article in the series, and the article on elections waits until elections actually run [18].

What to watch

  • Whether the next Squirix preview adds PreVote and RequestVote to the replication gRPC contract and has the host read the failover switches outside tests.
  • Whether the rewritten failover tests stop the key's owner and assert that a new leader took over, and when the delayed elections article appears.
  • The process changes the maintainer said the post would cover, and whether they make the build fail on MA0182 suppressions that cite unscheduled work.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories