BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Tailscale's 19 corrupt SQLite files show single-writer discipline is not corruption detection
Tailscale logged 19 corrupted shard databases in six months, and the first alarm came from a pipeline reading its S3 backups. Using SQLite correctly did not bound the failure.
The Engineer · Build desk
What happened
- Tailscale attributes months of control plane instability, from late last year into this one, to a single bug deep in SQLite that took months of forensics to isolate.
- It counted 19 separate database corruption incidents across six months before the underlying bug was resolved.
- Each repair required taking the affected shard offline, and the earliest windows ran over an hour before recovery was sped up.
- Established peer connections survived but went blind to changes; devices trying to come online could not join, and the console and API were unavailable.
- Tailscale posted global status page incidents each time, even though most shards and tailnets were never involved.
Why it matters
- constraint "We use SQLite the way SQLite says to use it" is no longer an available reliability argument: the documented single-writer pattern was in place throughout.
- exposure A shard could serve from a corrupt file until a backup consumer complained, and the tailnets on it had no way to know which shard they were sitting on.
- cost The bill fell on whichever customers happened to share a shard with a failing file, including re-entering configuration that did not survive recovery.
Nineteen corruptions in six months is about one every nine or ten days [7][17]. Tailscale calls SQLite corruption highly unusual and not something you should meet in normal operation [12], and that can be true at the same time as the rate above, because the unit differs. Per shard, the odds stayed small, and most shards were never involved at all [11]. Across the fleet, something broke every week and a half [20]. The status page reports the fleet [11].
The first signal, last August, did not come from the process that owned the database. It came from a data pipeline reading the S3 backups, and the integrity check was run against the backup copy [6]. That pipeline uploads the entire database file every few minutes [5], so a corrupt file reaches the bucket before anyone has looked at it, while the single Go process holding the write lock keeps serving [14]. Detection latency was therefore the snapshot interval plus whatever schedule the downstream reader ran on [18]. Nothing in the serving path was looking.
That is the part that generalises, more than the WAL internals. Tailscale had done what the documentation asks: one process, exclusive access, the pattern SQLite is built around [14], on a database it has run as primary storage since 2022 [4]. The fault was in SQLite itself, in WAL reset [3]. Correct usage bought a great deal and it did not bound corruption.
With the cause unknown, two mitigations were available and only one was reachable. Recovery got faster across the incident series [9]; incidence did not fall until the bug was identified months later [1][2][19]. That is a reasonable response to an undiagnosed fault, and it also means the team spent half a year treating corruption as a recurring operational event, paid for by stopping the shard's control plane each time [8].
For anyone running shard-per-SQLite, the practical reading is arithmetic. PRAGMA integrity_check exists and is cheap enough that Tailscale ran it by hand the moment it had a suspect file [6]. The number of database files scales with the shard count, and each one is a separate object that either has a named owner checking it on a cadence you set, or does not. Inheriting detection from a backup consumer means your mean time to detect is a property of your analytics schedule.
One more thing kept the blast radius survivable: these databases hold tailnet and device metadata, never private keys or traffic [15], so the worst outcome in the early incidents was a handful of recent devices and configuration changes that had to be re-entered [16]. Devices already connected stayed connected, blind to changes, while anything arriving during the repair window could not join [10]. A shard-per-SQLite design holding state that cannot be re-typed would have produced a very different post.
What to watch
- Whether the WAL-reset fix lands upstream in SQLite and in which release, since every other single-writer deployment inherits the same code path.
- Whether Tailscale wires scheduled integrity checks against live shard databases, and publishes the cadence, rather than relying on the backup reader.
- Whether status reporting moves to shard-scoped incidents now that global posts have been shown to overstate who was affected.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence66
- Adoption58
- Hype gap+14
- Incentives78
- Confidence62
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Tailscale's uptime was shaky at the end of last year and the instability continued into the new year; many of these outages were caused by a single bug deep in SQLite, and it took months of intense forensics to track down.
ReportedSupportedSource: Tailscale engineering blog2 sources— create a free account to open themView cited source - [2]
By summer, Tailscale says it is confident it has found the bug, understands it, and has fixed it.
ReportedSupportedSource: Tailscale engineering blog2 sources— create a free account to open themView cited source - [3]
Tailscale's post is titled "How Tailscale helped find the SQLite WAL-Reset bug" and describes the fault as a long-standing bug in the heart of the SQLite database that Tailscale helped uncover.
- [4]
Tailscale has used SQLite as its primary database since 2022, chosen because it is well known, reliable and widely used, which the company calls boring technology in a good way.
- [5]
Tailscale's backup pipeline takes a complete snapshot of each database every few minutes and uploads the entire SQLite file to an S3 bucket; it had run without incident since early 2023.
- [6]
In August last year a data pipeline reading the S3 backups reported an error in one database; Tailscale ran SQLite's PRAGMA integrity_check against the backup, found it corrupted, repaired the database and investigated the cause without success.
- [7]
Tailscale faced 19 separate instances of database corruption over six months before resolving the underlying bug.
- [8]
Whenever corruption occurred, Tailscale had to stop the control plane process on that shard while the database was repaired or restored, so tailnets on that shard lost their entire control plane during the recovery window.
- [9]
In the early incidents the downtime was over an hour; Tailscale gradually sped up the recovery process over subsequent incidents.
- [10]
A device joining a tailnet must fetch the device list from the control plane, so devices coming online during the SQLite downtime could not connect; already-connected devices stayed connected but could not learn about network changes, and affected tailnets temporarily lost the admin console and the Tailscale API.
- [11]
Tailscale posts a global incident on its status page even when only a small number of tailnets are affected, so many people saw status events for incidents that did not affect them; the majority of shards and tailnets were never involved in a corruption incident.
- [12]
Tailscale states that SQLite corruption is possible but highly unusual and not something you should encounter in normal operation.
- [13]
Clients reach the Tailscale control plane as a single public endpoint, controlplane.tailscale.com, but internally it is split into coordination servers called shards; each tailnet lives on one shard at a time and can migrate between them, and customers do not know which shard they are on.
- [14]
Each shard has an SQLite database holding all information about the tailnets on that shard, accessed exclusively by a single Go process that serves the control plane for those tailnets; Tailscale describes this single-writer design as exactly how SQLite is meant to be used.
- [15]
Because the control plane handles only configuration data, the affected databases contain metadata about tailnets and devices but never private encryption keys or network traffic.
- [16]
In the earliest incidents, recovery meant a handful of newly added devices or configuration changes did not persist and a small amount of metadata had to be re-entered.
- [17]
19 corruption incidents across six months averages roughly one every nine to ten days.
- [18]
On the evidence of the post, the detector of corruption sat outside the serving path: the writing process kept serving while a whole-file snapshot reached S3 and a downstream consumer read it, so time to detect was at least the snapshot interval plus the reader's schedule.
- [19]
Across the incident series the improvement was in recovery time rather than in incidence, because the root cause was not identified until months after the first corruption.
- [20]
The per-shard corruption rate stayed low while the fleet-wide rate was roughly one event every week and a half, and the status page reports the fleet rate rather than the per-shard rate.
Sources
1 independent publisher whose own reporting we read for this story.
- How Tailscale helped find the SQLite WAL-Reset bug
tailscale.com
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.