Build1 publisher3 min readPublished
Once You Have Added Leases And Fencing, Your Lock Is Just A Badly Timed Election
A dev.to series on distributed locking reaches the concession worth reading: exclusive ownership is rarely the problem, and hardening a lock rebuilds leader election with the expensive part kept.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A distributed lock is designed to coordinate ownership of a resource across multiple machines, ensuring that only one node can act on a shared resource at a time.
- Leases were introduced to handle abandoned locks caused by crashed or unreachable nodes; they make ownership temporary so it automatically expires if not renewed.
- Fencing tokens were added to address stale owners: even if a node believes it still holds a lock, a higher-numbered token can prevent it from making unsafe updates.
- According to the article, exclusive ownership is not always the problem being solved; in many systems the goal is not multiple machines competing for a resource but ensuring that one machine coordinates the rest.
- The article's example is a cluster of application instances running the same code, where every minute all of them check whether invoices should be generated; if all instances execute the scheduler independently, duplicate invoices become inevitable.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The third instalment of a dev.to series on distributed locking in practice arrives at the concession most teams reach late: exclusive ownership is not always the problem being solved, because in many systems the goal is simply that one machine coordinates the rest [4]. That matters because the standard hardening path for a lock, leases first and then fencing tokens, ends up reconstructing leader election while keeping the costly part [2][3].
Follow the progression as the piece lays it out. A distributed lock coordinates ownership of a resource across machines so that only one node acts on it at a time [1]. Crashed or unreachable holders leave locks abandoned, so ownership becomes a lease: temporary, and expiring unless renewed [2]. Stale owners, nodes that still believe they hold the lock, then get fenced, with a higher-numbered token blocking unsafe updates from the older holder [3]. Each addition is a correct response to a real failure mode. Collectively they change what you are holding. Expiring ownership, plus a monotonically increasing token, plus failover when the lease lapses, is the shape of leadership with an epoch number, and the residual difference is how often you re-run the handoff [14].
The worked example is one most scheduler teams have shipped: identical instances, each checking every minute whether invoices should be generated, with duplicate invoices inevitable if they all proceed [5]. The lock version is correct, and it pays a coordination round on every execution cycle [6]. At the stated one-minute cadence that is about 1,440 contended acquisitions per day [15], for a responsibility that does not move: one instance should own scheduling, and that assignment does not change frequently [7].
Leader election prices the same requirement differently. The cluster elects one coordinator, the others keep serving requests and stop competing unless a failure occurs [8]. The question shifts from who owns this operation right now to who is currently the leader of the cluster [10], and that leader can absorb the other work that wants a single owner anyway: cluster state, assigning work to nodes, health monitoring, shared configuration [9]. Election runs when leadership is lost, not per operation [12].
This is not a free safety upgrade, and the article does not sell it as one. Leaders crash, and when one does the cluster has no coordinator until a new leader is elected [11]. The exposure moves rather than disappears: from two nodes believing they hold the same lock, to an interval in which nobody is coordinating at all [17].
Which is where the published text runs out. The piece raises the right follow-up, how the system knows the leader has actually failed, and the text breaks off mid-sentence at "The honest answer is that it does" [13]. Failure detection is the load-bearing component of every design above, and it is the one left unwritten in part three of four [16].
Two things to watch. First, whether the concluding instalment puts numbers on detection: renewal intervals, timeout margins, and what the cluster does during the gap [13]. Second, a diagnostic you can run on your own system this week: if your lock acquisition rate is roughly constant and the winner is nearly always the same instance, you are holding an election every cycle and paying contention for a decision that is not actually in dispute [6][7].