Build1 publisher2 min readPublished
A 10-second inference probe pulled three of four healthy Macs out of a home LLM pool
One developer's 10-second inference probe pulled 3 of 4 healthy Macs from a Caddy pool where normal inference takes about 31.8 seconds. Stall detection now runs on real requests, so users who hit a wedged node supply the timeouts that get it removed.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- On a wedged node, the light /api/version query still answered in 0.25 seconds while inference requests to /api/chat never returned.
- The pool is four Macs running a 14-billion-parameter classification model, with Caddy's load balancer routing requests to whichever machines are alive.
- The replacement removes a node when real requests fail to return within 120 seconds, or come back 5xx, and that persists across the last 8 requests.
- Liveness went back to a periodic active hit on /api/version, leaving 'is it alive' and 'is the work stuck' to two separate monitors.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint On hardware where healthy inference takes about half a minute, fast active checks can only answer whether a process is up, since any inference probe must outwait normal latency to mean anything.
- cost Passive detection spends real user requests as its test signal, and each caller on a wedged node waits out a 120-second timeout until failures fill the 8-request window.
- decision Setting the timeout has to start from measured healthy latency; the 120-second limit sits near four times normal, giving up detection speed so slow requests can finish.
A 10-second cutoff is under a third of the 31.8 seconds a healthy node needs [1]. At that setting the probe is measuring latency, and a node running at its normal speed fails it. The author wrote that the probe cannot tell a node that is slow but normal from one that has stalled and will never return, so it drops healthy nodes [4].
The post describes these machines as doing light triage, fast [8]. On this hardware, fast is about 31.8 seconds [3]. That figure comes from a 14-billion-parameter model on Macs [7]. The lesson transfers where healthy inference is slower than the probe cutoff. The author still checks liveness actively with a periodic hit to a fast endpoint, and says the assumption breaks only when an active check is used to measure inference, "work that is slow even when healthy" [12].
The fault the probe was built to catch is real. The load balancer keeps rotating across every live machine, so only the requests that land on the wedged one vanish while the system as a whole still responds [13]. "The process is alive. But it is not doing its job," the author wrote [10].
The replacement uses the traffic that was already failing. A wedged /api/chat shows up on its own, because the real request times out [14]. The 120-second limit is nearly four times the normal response time [2], so a slow but healthy request has room to finish. The failure has to persist across the last 8 requests, so a single bad request does not pull a node [5].
Users pay for this design. Each timeout on a wedged node is a real caller who waited 120 seconds, and the node stays in rotation until the failures persist across the window [4]. The rule also needs traffic to see anything. A node that gets no requests records no timeouts [5]. The author saw wedges when several heavy requests piled onto one node, and when a request hit a node idle long enough that the model had to be loaded back into memory [11].
The post does not report false-removal counts or a measured detection time for the passive rule, so whether it avoids false removals is argued from the design and not measured. I think the split is right for this pool. A 0.25-second hit on /api/version answers whether the process is up, and real requests answer whether the work comes back [6]. The author made the change and reverted it the same day [1].
What to watch
- Published false-removal counts or detection times for the 120-second, 8-request rule would show whether it avoids pulling healthy nodes in practice.
- Cold-start reloads were one wedge trigger; if a reload on a healthy idle node ever exceeds 120 seconds, the passive rule would count it as a failure.