Build1 publisher3 min readPublished
A retry ID built from fixed state kept a trading bot's stop-loss repair failing for hours
A paper-trading bot's stop-loss self-heal failed on every retry for hours while its health check passed all 55 checks. Each recovery step was correct on its own, and the fault only appeared when the whole chain met the broker's uniqueness rule.
The Engineer · Build desk
What happened
- The recovery chain ran from a transient close failure through a handler, a health check and a self-heal order, and the author judged each of the four steps correct on review.
- The author found the fault by running the sync command by hand and reading the broker's raw error, not by reading the code.
- The fix appends eight hex characters of a fresh UUID to each attempt's ID so no two attempts send the same string.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Line-by-line review cannot catch this class of fault because the breaking rule, the broker's uniqueness check, lives in no function under review.
- exposure Any retry against an external system that treats an ID as unique, including idempotency keys, filenames and job names, can loop forever if the ID comes only from fixed state.
- decision A green structural health check cannot stand in for recovery monitoring; the self-heal path needs its own consecutive-failure counter and a known-safe fallback.
The rule that broke the chain belongs to the broker. Every client_order_id must be globally unique [7]. The self-heal built its stop order ID by appending "_stop" to the position's original ID, a string made entirely from state that does not change [6]. One transient failure on the first attempt is enough to spend that ID at the broker [8]. Every later attempt sends the same string and gets the same rejection [8]. The loop "would have kept failing until someone changed the code, no matter how many times it ran," the author wrote [9].
Review could not catch this. Each function did what it claimed, and the author had read all of them more than once [4]. "Reviewing this code teaches you nothing, because there's nothing wrong with it," the author wrote [5]. The defect exists only where a deterministic ID meets a retry against a system that enforces uniqueness outside the repository [6][7].
The health check ran seven layers of verification and passed 55 of 55 checks through the whole incident [11]. It confirmed that files parsed, agents started, credentials loaded and cron jobs existed [11]. "It was right about all of that. The system was structurally intact. It just wasn't correct," the author wrote [12]. The author splits the question in two. Structural health asks whether the machine is running. Logical correctness asks whether what the logs believe matches what reality shows [13]. Only the first had an instrument. The self-heal itself decides from a local field, placing a stop when stop_order_id is missing [4]. Nobody read the broker's answer until the author ran the sync command by hand and read the raw error [14].
The fix appends eight hex characters of a fresh UUID to each attempt's ID [10]. The author found the bad ID by grepping for string-concatenated IDs [15]. It is a good fix for the collision, and it is one line. It also means the broker can no longer tell a retry from a new order. If an attempt ever lands at the broker and only its response is lost, the next attempt places a second stop on the same position. I'd pair the suffix with a read of the broker's open orders before placing anything. A local field is exactly the kind of belief the author's own correctness question says to check against reality [13].
The whole-chain test the author prescribes is to trigger the recovery path against a real sandbox or a realistic mocked failure, then read what actually comes back [16]. "Reading a recovery path is not testing it," the author wrote [16]. The alerting piece already existed in the codebase for a different risky action. It counts consecutive failures and, past a threshold, stops trying and falls back to a known-safe state [17]. According to the author, it never spread to the stop-loss retry on its own [17]. While the retry spun, "nothing anywhere told me so," the author wrote [3].
What to watch
- Whether the author publishes the circuit-breaker design for the self-heal path, including what the known-safe fallback is for an unprotected position.
- Whether the self-heal starts reading the broker's open orders before placing a stop; without that check the per-attempt suffix leaves duplicate stops possible after a lost response.
- Whether the grep for string-concatenated IDs turns up other retry paths built only from fixed state.