Skip to content

Build1 publisher3 min readPublished

An unreviewed stratum-1 change let one GPS card send Telstra's network back to 2006

Telstra's external audit traces the July 8 outage that hit nearly nine million customers to one unpatched GPS card holding stratum 1. Operators who change stratum assignments without review can have one bad receiver set every downstream clock.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying An unreviewed stratum-1 change let one GPS card send Telstra's network back to 2006
Photo: abc.net.au

What happened

  • A power-supply swap in Melbourne rebooted the card, and firmware missing the GPS week-rollover fix brought it up 1,024 weeks in the past, in November 2006.
  • Nine months earlier the card had been switched on to stop a flapping time source, making Melbourne stratum 1, level with Australia's national time source, with no review.
  • The outage hit 45% of all calls and data sessions that day, put errors on 604 Triple Zero calls, and took down trains and card terminals.
  • The vendor's firmware fix had appeared in bulletins in 2022 and January 2026, according to the Guardian's Senate coverage, and was never installed.
  • Telstra has since moved all three sites off the old time servers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A fine of up to $30 million would be about 1,000 times the roughly $30,000 Telstra's CEO told the Senate it costs to replace the 2011 chassis.
  • exposure Receivers that rode through the April 2019 rollover by staying powered, as this card did, still face the wrong-era guess at their next cold start if their firmware predates the fix.
  • decision With the engineers who knew the change on mandatory rest, recovery needed a configuration record TAP found did not exist, so time-source config has to be written where the on-call shift can read it.

NTP sorts clocks into strata [23]. Stratum 0 is a reference clock, such as an atomic clock or a GPS receiver. A server attached to one is stratum 1, and anything syncing from that server is stratum 2 [23]. A server with several sources has two defences. Lower stratum carries more weight, and a source that disagrees with the rest is treated as a false ticker and voted out [5].

Neither defence looks at the calendar. "NTP does not ask whether a date is plausible; it asks whether a source disagrees with the others," Sven-Christian Ebenhag wrote in an analysis for Netnod, which runs Sweden's national time distribution [12][21]. On July 8, according to a dev.to walk-through of the audit, the broken card outranked every other clock in the tree, and any clock that could have contradicted it was itself synced to Melbourne [5]. "The protocol worked. The architecture did not," Ebenhag wrote [13].

Part of that architecture dates from 2020, when a switch to peering made timing loops possible [14]. In ntpd the difference is one word per line. The walk-through gives an illustrative config and says it is not Telstra's [22]:

``` # client/server: take time from above, never give it back server ntp1.upstream.example iburst # peering: two servers at the same level may sync from each other peer ntp-mel.internal.example ```

Peering is meant for servers of equal rank to back each other up. After the 2020 cross-wiring, one peer had become the only real source for the other [14]. "Two servers fed by the same GNSS receiver is still one source, but counted twice," Ebenhag wrote [15].

The wrong date came from a 10-bit counter. GPS broadcasts the week number in ten bits, so the count tops out at 1,023 and wraps every 1,024 weeks [16]. Each era of 1,024 weeks is 7,168 days, about 19.6 years [1]. The wraps so far fell in August 1999 and April 2019 [16]. A receiver running through a wrap keeps counting. One that cold-starts has to guess its era from firmware, and older firmware guesses an older era [17].

I think the outage needed all three faults at once: the old firmware, the unreviewed promotion and the 2020 peering layout [3][4][14]. The promotion is where I would start. A stratum number decides which box every other box believes. In my view, stratum assignments and peer lines belong in the same change review as routing policy. I would also want at least one cross-check fed by a receiver that shares nothing with the primary. NTP will not reject a 2006 timestamp on its own [12], so a monitor outside the protocol should alarm when served time jumps by years.

The published account of Telstra's fix does not describe how the new servers are ranked or peered [8]. The 10-page TAP report is still the postmortem I would hand a new on-call engineer [9]. "It was an outage that should not have happened," Telstra CEO Vicki Brady wrote in her summary post [10].

What to watch

  • Whether the regulator imposes the possible fine of up to $30 million, and at what level.
  • Whether Telstra publishes how its replacement time servers are ranked and peered, and which independent source they cross-check against.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories