Skip to content

Build1 publisher3 min readPublished

A supervisor rebuilt from spawn_link and receive shows where restart bugs actually live

A dev.to walkthrough reconstructs one-for-one supervision from BEAM primitives. The exercise separates the mechanism, links and exit signals, from the policy that decides what gets restarted.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A supervisor rebuilt from spawn_link and receive shows where restart bugs actually live
Generated illustration

What happened

  • Part 5 of the dev.to series 'Building Distributed Systems in Elixir' builds a small supervisor from scratch using spawn/1, spawn_link/1, Process.flag(:trap_exit, true), send/2 and receive/1.
  • The exercise uses no GenServer and no OTP Supervisor.
  • The author states the goal is not to replace OTP but to understand the mechanism and policy that an OTP supervisor provides.
  • A linked worker that crashes sends an exit signal to the process linked to it; by default that failure propagates and can terminate both processes.
  • Once exit trapping is enabled with Process.flag(:trap_exit, true), an incoming exit signal becomes a mailbox message of the form {:EXIT, pid, reason}.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Part 5 of a dev.to series on building distributed systems in Elixir rebuilds a supervisor by hand, with no GenServer and no OTP Supervisor, using only spawn/1, spawn_link/1, Process.flag(:trap_exit, true), send/2 and receive/1 [s1c1][s1c2]. The author is explicit that the point is not to replace OTP but to see the mechanism and the policy that an OTP supervisor provides [s1c3], and that division is the one worth keeping in your head the next time a node is churning through restarts.

The mechanism half is small. By default, a linked worker that crashes sends an exit signal that propagates and can terminate both processes [s1c4]. Set the trap_exit flag and the same signal arrives as an ordinary mailbox message, {:EXIT, pid, reason} [s1c5]. That is detection, and detection is only the beginning [s1c6]. Messages sent to a dead local PID do not bring the process back [s1c7], so something else has to decide what happens next.

The alternative the article rules out is polling with Process.alive?, for four reasons: detection is delayed until the next check, restart logic ends up scattered through the application, every worker has to be checked repeatedly, and a liveness check does not tell you why the worker stopped [s1c8]. Links and exit signals are already event-driven failure detection on the BEAM [s1c9].

The policy half is a list of responsibilities: trap exits, start and link both workers, record which PID maps to which logical worker name, wait for messages, restart a worker after an abnormal exit, leave healthy siblings running, and do not restart after a normal exit [s1c10]. The article calls that a minimal one-for-one restart strategy [s1c11].

The worker makes the policy input concrete. Worker.start/1 uses spawn_link/1, which starts the process and creates the link atomically [s1c12]. The loop carries a name and a processed-job count, replying {:done, name, value, new_count} to each job [s1c13]. A :crash message raises, which terminates the worker with an abnormal exit reason [s1c14]. A :stop message returns :ok, letting the function finish with reason :normal [s1c15]. The article states that this distinction determines whether the supervisor restarts the worker [s1c16]. In this implementation the exit reason is the whole of the policy input [s1d1], which is why a worker that exits cleanly will never come back and a worker that raises always will [s1d2].

One detail is easy to skim past: the supervisor itself is created with spawn/1 rather than spawn_link/1, because the caller running the demonstration is not part of the supervision relationship [s1c17]. Get that wrong in real code and your test harness becomes a supervision parent.

Worth noting what the published excerpt does not reach: it stops mid-sentence as the supervisor enters init/1 [s1c18], and the enumerated responsibilities contain no restart budget at all [s1d3]. A hand-rolled one-for-one loop with no intensity limit will restart a permanently broken child forever, which is exactly the failure mode teams misattribute to OTP.

Watch whether later parts of the series introduce restart intensity limits and strategies beyond one-for-one; the previous part covered process links [s1c19], so the ladder is being climbed in order.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories