Build1 publisher3 min readPublished
The network already knew: UPS and WAN state as keys a cluster can reconcile against
A homelab operator wired UniFi WAN failover and UPS battery state into Kubernetes as declarative state keys, so optional workloads scale themselves to zero while the condition holds.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- UniFi Reactor is a Kubernetes operator, built by Robbe Verhelst, that polls the UniFi Network API, normalizes what it sees into state keys, and reconciles Kubernetes Automation resources against those keys. Docs at reactor.robbeverhelst.com.
- The author's UniFi gear knew when the WAN failed over and when the UPS switched to battery; Kubernetes did not. He writes that the network already had the state and the cluster just needed to react to it.
- If the main WAN failed at 3 AM, qBittorrent could keep seeding over a metered backup link.
- If power dropped, the UPS could be counting down its remaining runtime while the cluster continued background ML jobs, backups and other work that did not need to happen during an outage. The author added a UPS and backup internet to the homelab a few weeks before writing.
- Example state keys: wan (primary or backup); internet (ok, degraded, down); ups (online or on-battery); ups.battery (normal, low, critical); devices (all-online or degraded); device.<name> (online or offline).
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Robbe Verhelst has published UniFi Reactor, a Kubernetes operator that polls the UniFi Network API, normalizes what it sees into state keys, and reconciles Automation custom resources against those keys [1]. The gear already held the facts that matter during an outage; what was missing was a representation the scheduler could act on [2].
The failure mode he describes is specific enough to repeat. His UniFi equipment knew when the WAN failed over and when the UPS switched to battery, and Kubernetes did not [2], so a primary WAN failure at 3 AM left qBittorrent seeding over a metered backup link [3], and a power cut left the UPS counting down its remaining runtime while the cluster carried on with background ML jobs and backups [4].
The keys are deliberately coarse: wan primary or backup; internet ok, degraded or down; ups online or on-battery; ups.battery normal, low or critical; devices all-online or degraded; and per-device online or offline [5]. Six families, total [6]. An Automation matches on a provider plus a state map, applies its actions while the match holds, and applies onExit actions when it stops holding [7]. The first rule he actually deployed scales the immich-machine-learning Deployment to zero when ups reads on-battery and back to one replica when mains power returns [8]. He calls it boring in the right way, on the grounds that nobody minds if photo indexing pauses during a power cut [9].
The design choice that carries the weight is state over events: polling is the source of truth, and webhooks may become a fast path later but should not be the mechanism of record [10]. His reasoning is that one-shot events are easy to miss, because controllers restart, networks flap and webhooks fail, and an edge-triggered system can end up stranded in the wrong mode [11]. That is the transferable part. An alert says a condition occurred; a state key says a condition currently holds, and "currently holds" is the only thing a reconciler can act on. Verhelst is explicit that monitoring tells you about state while Reactor changes cluster behaviour while that state holds [12], and that for notification-only cases Prometheus is probably still the better tool [13]. He already runs Prometheus, Grafana and Gatus [14].
The honest limits are in the post. The flagship case, pausing downloads while on backup WAN, exists as a YAML example that stops qBittorrent and restores it when the primary link returns [15], but he frames it as what he wants once failover is verified end to end [16]. The follow-ons he lists, suspending offsite backup CronJobs on metered WAN, scaling down Jellyfin remote streaming, disabling guest WiFi during failover, and notifying when the primary link is up but the internet is still down, are not all implemented [17]. The API group is v1alpha1 [18]. And the restore path carries a literal replicas: 1 rather than the pre-outage count [19], which is fine for a single-replica Deployment and wrong the moment an autoscaler owns that field.
Worth watching: whether onExit learns to capture prior state instead of asserting a constant; whether the low and critical battery tiers get distinct policies rather than one binary shed [5]; and what the poll interval and any flap hold-off are, since the post does not state them [20].