Build1 publisher3 min readPublished
A one-millisecond flaw in legacy NATS code stopped UK departures for four and a half hours
NATS traced a four-and-a-half-hour halt to UK departures to a legacy defect that corrupted a squawk-code update inside a one-millisecond window. Its first symptom, a link drop that healed in 45 seconds, was logged as having no ongoing operational impact.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- At 10:00 a correctly made squawk-code request was paused mid-update for a higher-priority message and resumed with corrupted output.
- The corrupted data made London Area Control's system time out, and it cut its link to the national flight data system, as designed, to protect both.
- Clearing the one bad record in London's airspace took a controlled restart of the national system, which meant restrictions across the whole UK.
- NATS had forecast about 8,000 flights that day and handled 6,094.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint While clearing a corrupted NAS record means restarting the NAS, a fault in one sector's traffic becomes restrictions at every airport and centre the system feeds.
- decision Teams running coupled systems have to choose whether a protective disconnect that heals itself closes an event or opens one; here the event was closed while the corrupt record stayed in the NAS.
- cost One record cost the day about 1,906 flights against forecast, close to a quarter of planned traffic.
- exposure A window this narrow had never fired before, so a clean operating history is weak evidence that legacy priority-scheduled code is free of similar half-finished updates.
The NAS switches between jobs by priority, and the report treats that as routine: "Switching between different activities in response to prioritised requests is a normal function of the system" [13]. The defect needed a higher-priority request to arrive "during that exact millisecond while the original request was part-way through updating a value" [22]. The interrupted request was clean. It was "made correctly and there was nothing abnormal or invalid about the associated flight plan," the report says [14]. "Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally," according to the report [5].
The report never uses the words "race condition". Programmers on Hacker News used them at once [17]. "Looks like a race condition (with a 1ms window) during squawk allocation," macguillicuddy wrote [18]. Another commenter, crote, wrote: "something which was supposed to be an atomic operation was split into two by the preempting task" [19]. The NATS code is not public, and the report is silent on exactly what failed when the interrupted request resumed [20].
The dev.to post's author offers a generic sketch of the bug class and labels it, in capitals, NOT the NAS code [21]. It sets `record->status = ALLOCATING;` and then `record->code = pick_free_code();`, with a comment that a higher-priority task can run between the two lines [21]. Anything that reads the record in that gap sees the new status next to the old code [21].
Isolating on a timeout is the right default for a link between two safety systems, and London Area Control's side did it as designed [6]. The 10:02 log entry was accurate about the link [7]. The corrupted record stayed in the NAS until a controlled restart of the national system cleared it [8]. NATS's first public post, saying it was "investigating a technical issue", went out at 14:16 BST [15]. That was 4 hours and 14 minutes after the first drop [2].
Where one record can force a national restart, I'd classify a self-healed protective disconnect by what triggered it. Whether the link came back is a separate question. A timeout caused by corrupted data from the core system is still an open incident after the link returns, because the data has not been repaired [8]. The report found no operator error and no sign of an attack [10]. If the people on shift did their jobs, the weakness sits in the classification they had to work with.
The coupling sets the scale. NAS data feeds systems at several UK airports and other control centres [11]. London Area Control alone handles about 6,500 flights on a typical September day [12]. More than 2,000 flights were delayed, cancelled or diverted [2], and the backlog took more than two days to clear, with "hundreds of thousands of passengers' travel plans disrupted", according to the NATS press release [16]. NATS says the permanent software fix is written and being tested [23].
What to watch
- Whether the final NATS report explains what went wrong on resume and what the permanent fix changes in the squawk allocation code.
- Whether NATS changes how self-healed links between London Area Control and the NAS are classified and escalated.
- Whether NATS gains a way to clear a single corrupted NAS record without a national restart.