Build1 publisherNot yet confirmed elsewhere2 min readPublished
Firebase's zero-length flag entry crashed iOS apps at launch for up to six hours
Firebase pushed a backend change on 29 September that crashed every iOS app running its SDK with analytics on, some for up to six hours. Phones kept the bad response cached after Google's rollback, so apps that start the SDK at launch need an off switch Firebase cannot override.
The Engineer · Build desk

What happened
- Firebase's later postmortem says the team was alerted 20 minutes after the change rolled out, through crash alerts and GitHub issues.
- An outside developer traced the crashes to a zero-length entry that was breaking the SDK and posted the root cause on GitHub at 6:37pm PDT, before any Google engineer had acknowledged the outage.
- A Firebase engineer first acknowledged the outage on the developers' GitHub ticket at 6:51pm PDT, an hour and ten minutes after the crashes began.
- Firebase began rolling back the offending backend change at 7:24pm PDT and finished at 8:16pm PDT.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint When a vendor pushes config to devices, customers stay down for as long as the client cache holds the bad copy, so the vendor's rollback-complete time understates how long they were down.
- exposure A setting decided who was hit: apps with Firebase analytics enabled. Each vendor feature an app switches on at launch widens what one vendor backend push can break.
- cost The 50 minutes between Firebase's internal alert and its first public word fell on app teams, whose only shared channel in that window was a GitHub thread they had opened themselves.
The failure came in through configuration. Firebase delivers flags from its backend to the SDK inside each app, so the change that broke iOS clients was made on the backend [9]. Developers on the ticket, which passed 100 comments, asked Google to roll it back [8][9].
Rolling back the server did not reach phones that had already stored the bad response. Those apps kept serving the cached copy for about four more hours after the rollback, according to The Pragmatic Engineer [12]. The rollback itself took 52 minutes by the ticket's timestamps [20].
Gergely, who writes the Pragmatic Engineer newsletter [16], wrote that Firebase showed "shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents" [15]. His timeline supports him on diagnosis speed. The outside developer posted the root cause 56 minutes into the outage [17]. That was 14 minutes ahead of Google's first word on the ticket [18]. He asked Google whether the ticket or its own monitoring raised the alarm, and had no reply when he wrote [7].
Google coined the term Site Reliability Engineer and wrote the SRE book [14]. Firebase's status page showed all systems green during the outage and after it [6]. The team closed the incident with a short report saying the outage was resolved [13].
I think the lesson for app teams follows from how this failure travelled. The bad state arrived as vendor config and then outlived the vendor's fix in a cache on the device [9][12]. A switch that helps has to sit in front of both. It also has to work when the app cannot stay open long enough to fetch anything, since these apps crashed on first opening [1]. On each launch, in order:
1. Check for a marker the app writes just before initializing the SDK and clears once start-up succeeds. If the marker is still there, the last launch died inside the SDK, so skip it this time. 2. Otherwise read a flag your own service controls from local storage, and initialize the SDK only if it is on. 3. After launch, refresh that flag from your own endpoint with a short cache lifetime.
The first step needs no network. A device that crashed once would skip the SDK on its next launch. The gate only works if the SDK stays idle until the app calls it. The newsletter does not report whether any affected team had a switch like this, or how one fared.
What to watch
- A fuller Firebase postmortem that explains why public acknowledgement trailed the internal alert, and whether monitoring or the GitHub ticket caught the crash first.
- An iOS SDK release that tolerates a zero-length flag entry; that fix would close this specific path on devices that update.
- Any change to how long the Firebase SDK caches server responses, since that cache set the length of this outage for many devices.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence58
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The outage started on Tuesday 29 September at 5:41pm PDT, when iOS apps using the Firebase SDK started to crash upon first opening.
ReportedSupportedSource: The Pragmatic Engineer2 sources— create a free account to open themView cited source - [2]
Every iOS app that uses the Firebase SDK with analytics enabled was affected.
ReportedSupportedSource: The Pragmatic Engineer2 sources— create a free account to open themView cited source - [3]
Before a Google engineer acknowledged the incident, an external developer found the root cause at 6:37pm PDT: a zero-length entry that was crashing the SDK.
ReportedSupportedSource: The Pragmatic Engineer, citing GitHub2 sources— create a free account to open themView cited source - [4]
In the postmortem later published by the Firebase team, they wrote that the team was alerted 20 minutes after the rollout, via crash alerts and GitHub issues.
ReportedSupportedSource: Firebase postmortem, as reported by The Pragmatic Engineer2 sources— create a free account to open themView cited source - [5]
Gergely asked Google/Firebase whether the team was alerted via the GitHub ticket or by Google's own monitoring and had not had a response at the time of writing.
ReportedSupportedSource: The Pragmatic Engineer2 sources— create a free account to open themView cited source - [6]
The official Firebase status page showed all systems green during and after the outage and was never updated to indicate it.
ReportedSupportedSource: The Pragmatic Engineer2 sources— create a free account to open themView cited source - [7]
Gergely asked Google/Firebase whether the alert came via the GitHub ticket or Google's own monitoring, and had no response when he wrote.
ReportedSupportedSource: The Pragmatic Engineer2 sources— create a free account to open themView cited source - [8]
Developers of affected apps opened a GitHub ticket, which drew 100+ comments; 'it's crashing for me too!' was oft-repeated.
- [9]
The flags are shipped by the backend, so the offending change was a backend one; the community urged Google to roll it back.
- [10]
At 6:51pm PDT, an hour and ten minutes after the crashes started, an engineer on the Firebase team acknowledged on GitHub that they were aware of the outage.
- [11]
The Firebase team started rolling back the offending backend change at 7:24pm PDT, and the rollback completed at 8:16pm PDT.
- [12]
Apps that had cached the incorrect server response served this cache for an additional four hours, and so kept crashing for up to six hours.
- [13]
The Firebase team closed the outage with a short report effectively saying there had been an outage and it had been resolved.
- [14]
Google coined the term 'Site Reliability Engineer' and wrote the SRE book.
- [15]
shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents
- [16]
Gergely writes the Pragmatic Engineer Newsletter.
- [17]
The external developer posted the root cause 56 minutes after the crashes started.
- [18]
The root cause was posted 14 minutes before a Firebase engineer acknowledged the outage.
- [19]
Public acknowledgement came 50 minutes after the internal alert the postmortem describes.
- [20]
The rollback took 52 minutes from start to completion.
Sources
1 independent publisher whose own reporting we read for this story.
- blog.pragmaticengineer.comThe Pulse: Firebase’s global outage & poor response
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.