Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Tailscale's 19 corrupt SQLite files show single-writer discipline is not corruption detection

Tailscale logged 19 corrupted shard databases in six months, and the first alarm came from a pipeline reading its S3 backups. Using SQLite correctly did not bound the failure.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Tailscale attributes months of control plane instability, from late last year into this one, to a single bug deep in SQLite that took months of forensics to isolate.
  • It counted 19 separate database corruption incidents across six months before the underlying bug was resolved.
  • Each repair required taking the affected shard offline, and the earliest windows ran over an hour before recovery was sped up.
  • Established peer connections survived but went blind to changes; devices trying to come online could not join, and the console and API were unavailable.
  • Tailscale posted global status page incidents each time, even though most shards and tailnets were never involved.

Why it matters

  • constraint "We use SQLite the way SQLite says to use it" is no longer an available reliability argument: the documented single-writer pattern was in place throughout.
  • exposure A shard could serve from a corrupt file until a backup consumer complained, and the tailnets on it had no way to know which shard they were sitting on.
  • cost The bill fell on whichever customers happened to share a shard with a failing file, including re-entering configuration that did not survive recovery.

Nineteen corruptions in six months is about one every nine or ten days [7][17]. Tailscale calls SQLite corruption highly unusual and not something you should meet in normal operation [12], and that can be true at the same time as the rate above, because the unit differs. Per shard, the odds stayed small, and most shards were never involved at all [11]. Across the fleet, something broke every week and a half [20]. The status page reports the fleet [11].

The first signal, last August, did not come from the process that owned the database. It came from a data pipeline reading the S3 backups, and the integrity check was run against the backup copy [6]. That pipeline uploads the entire database file every few minutes [5], so a corrupt file reaches the bucket before anyone has looked at it, while the single Go process holding the write lock keeps serving [14]. Detection latency was therefore the snapshot interval plus whatever schedule the downstream reader ran on [18]. Nothing in the serving path was looking.

That is the part that generalises, more than the WAL internals. Tailscale had done what the documentation asks: one process, exclusive access, the pattern SQLite is built around [14], on a database it has run as primary storage since 2022 [4]. The fault was in SQLite itself, in WAL reset [3]. Correct usage bought a great deal and it did not bound corruption.

With the cause unknown, two mitigations were available and only one was reachable. Recovery got faster across the incident series [9]; incidence did not fall until the bug was identified months later [1][2][19]. That is a reasonable response to an undiagnosed fault, and it also means the team spent half a year treating corruption as a recurring operational event, paid for by stopping the shard's control plane each time [8].

For anyone running shard-per-SQLite, the practical reading is arithmetic. PRAGMA integrity_check exists and is cheap enough that Tailscale ran it by hand the moment it had a suspect file [6]. The number of database files scales with the shard count, and each one is a separate object that either has a named owner checking it on a cadence you set, or does not. Inheriting detection from a backup consumer means your mean time to detect is a property of your analytics schedule.

One more thing kept the blast radius survivable: these databases hold tailnet and device metadata, never private keys or traffic [15], so the worst outcome in the early incidents was a handful of recent devices and configuration changes that had to be re-entered [16]. Devices already connected stayed connected, blind to changes, while anything arriving during the repair window could not join [10]. A shard-per-SQLite design holding state that cannot be re-typed would have produced a very different post.

What to watch

  • Whether the WAL-reset fix lands upstream in SQLite and in which release, since every other single-writer deployment inherits the same code path.
  • Whether Tailscale wires scheduled integrity checks against live shard databases, and publishes the cadence, rather than relying on the backup reader.
  • Whether status reporting moves to shard-scoped incidents now that global posts have been shown to overstate who was affected.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence66
Adoption58
Hype gap+14
Incentives78
Confidence62
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Tailscale's uptime was shaky at the end of last year and the instability continued into the new year; many of these outages were caused by a single bug deep in SQLite, and it took months of intense forensics to track down.

    ReportedSupportedSource: Tailscale engineering blog2 sources— create a free account to open themView cited source
  2. [2]

    By summer, Tailscale says it is confident it has found the bug, understands it, and has fixed it.

    ReportedSupportedSource: Tailscale engineering blog2 sources— create a free account to open themView cited source
  3. [3]

    Tailscale's post is titled "How Tailscale helped find the SQLite WAL-Reset bug" and describes the fault as a long-standing bug in the heart of the SQLite database that Tailscale helped uncover.

Sources

1 independent publisher whose own reporting we read for this story.

  1. tailscale.com

    1 article · August 23, 2026

    How Tailscale helped find the SQLite WAL-Reset bug

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories