Skip to content

Build1 publisher3 min readPublished

A three-Pi HA control plane that ended up less reliable than the one node it replaced

A homelab k3s build put embedded etcd on microSD cards and USB Ethernet. The reported latency numbers point at I/O contention, not the consensus timeouts the write-up blames.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying A three-Pi HA control plane that ended up less reliable than the one node it replaced
Photo: mei-home.net

What happened

  • A homelab author publishing on dev.to as dwoitzik planned and built three Raspberry Pi 5 control-plane nodes running k3s with embedded etcd, replacing three Proxmox VMs.
  • After three weeks of monitoring, the author concluded the k3s-on-Pi cluster was less reliable than the single-node setup it replaced, and reverted.
  • The author describes the failure as undramatic: no kernel panic and no cluster death, but a slow accumulation of fragility.
  • Cluster spec: 3x Raspberry Pi 5 (8GB), 256GB A2-rated microSD cards, gigabit Ethernet via USB 3.0 adapter, k3s v1.31 with embedded etcd.
  • The post states gigabit Ethernet was provided via USB 3.0 adapter because the Pi 5's native Ethernet is limited; it gives no measurement or further detail for that characterisation.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A homelab operator publishing as dwoitzik on dev.to replaced three Proxmox VMs with three Raspberry Pi 5 nodes running k3s control plane and embedded etcd, then reverted after three weeks because the cluster was less reliable than the single-node setup it had replaced [1][2]. Nothing crashed, which is what makes the write-up worth reading: the fragility sat in the storage and the interconnect underneath the consensus layer, not in Kubernetes [3].

The build was three Pi 5s with 8GB of RAM, 256GB A2-rated microSD cards, gigabit Ethernet through USB 3.0 adapters, and k3s v1.31 with embedded etcd [4]. The adapters were chosen because the author considers the Pi 5's native Ethernet limited, a description the post does not quantify [5]. The post also carries Amazon affiliate links for hardware the author says he owns and uses [6].

The write pattern is the core of the argument, and it is sound. Every API operation, from pod scheduling to configmap updates to secret rotation, becomes an etcd write, and embedded etcd writes continuously to the local filesystem [7]. The A2 rating of roughly 150 MB/s sequential is the wrong number for that workload; the random small-block writes are a different story [8]. What the author measured was etcd response times occasionally jumping from 10ms to 500ms or more with no CPU load, network congestion or memory pressure to explain it [9], against microseconds for the same writes on the NVMe-backed VMs [10].

Where the post overreaches is the causal chain to outages. It attributes brief API unavailability to etcd leader elections triggered when write latency exceeds the default 5s election timeout [11], with a re-elected leader on an equally slow card failing the same way, so the API server drops out while every node reports healthy [12]. The arithmetic does not reach that far. A 500ms spike is one tenth of the 5s budget [1]. The network penalty the post computes, about 0.5ms per USB adapter hop and roughly 2.5ms for a full write-replicate-acknowledge round trip [13] versus about 0.1ms on the VMs, which it calls a 25x increase [14][2], is about 0.05 percent of that same budget [3]. Twenty-five times a very small number is still very small. The post reports no leader-change counts, no outage durations, and no etcd fsync or backend commit figures beyond the 10ms-to-500ms range [15].

The co-tenancy mechanism holds up better. The Pis were already running AdGuard Home and Unbound for network DNS, Keepalived for a VIP, and node_exporter, so k3s made four services contending for one CPU, one memory pool and one card [16]. The author describes AdGuard cache refreshes landing at the same time as etcd compaction, both hitting the card at once [17]. That is queueing, and it does not need a five-second stall to be damaging: an API server that intermittently answers 50 times slower is already a poor dependency [4].

The revert, logged as ADR-014, put the sole control plane and etcd on one NVMe-backed VM with two agent-only workers [18]. The stated cost is that losing the server loses the cluster, with no failover [19], which in this context means minutes of restart rather than an outage with users attached [20]. The Pis kept DNS and Keepalived, work the author characterises as low-write and mostly reads after cache warm-up [21].

If you run the same layout, instrument etcd disk fsync and backend commit latency and count leader changes before you tear anything down. If fsync dominates, moving etcd's data directory off the shared card is cheaper than surrendering HA; if leader changes are rare, the outages had a different cause. Three of anything buys availability, not reliability.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories