Product1 distinct publisher3 min readPublished
Both figures come from the providers' own status pages, and the sharper test is the afternoon of September 3, when OpenAI and Anthropic were degraded inside the same window, which is what a two-subscription fallback has to survive.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The reflex, when a model starts returning errors, is to doubt your own machine and switch browsers or sign out and back in [20]. The answer is usually published: OpenAI at status.openai.com, Anthropic at status.claude.com, xAI at status.x.ai, with Gemini in the Google Workspace Status Dashboard [21]. Those pages lag, though. On September 3 the Downdetector reports spiked before the official entries appeared, with ChatGPT peaking above 35,000 in the US [19] [7].
Ninety days is 129,600 minutes [1]. OpenAI's 99.64 percent leaves 467 of them unavailable, just under eight hours, and Anthropic's 99.4 percent for claude.ai leaves 778, a little over thirteen [11] [12]. The figure sitting beside OpenAI's changes how you read it: the APIs run at 99.94 percent and Codex at 100 [8]. Over the same window 99.94 percent is about 78 minutes, so the consumer app was out roughly six times as long as the interface developers build against [2] [3].
That gap matters less than what the afternoon of September 3 did to the arithmetic of diversification. If ChatGPT and claude.ai failed independently, the joint probability would be 0.36 percent times 0.6 percent, which over 90 days works out to 2.8 minutes of simultaneous downtime [4]. OpenAI's incident ran 14:58 to 16:55 UTC and Anthropic's second ran 13:26 to 16:16 [2] [4], an overlap of 78 minutes [5], about 28 times the quarter's independent allowance in a single afternoon [6]. Anthropic logged a shorter Claude Sonnet 5 incident from 12:37 to 12:56 UTC as well [3]. To keep this honest: "elevated errors" is degradation rather than a blackout, and OpenAI logged 60 incidents between July 1 and September 3, roughly one a day [9], almost none of which anyone outside noticed.
Neither company has given a cause in its incident notes [16]. Much of the coverage pointed at Azure East US, where ChatGPT, Claude and Grok are said to draw compute [14], but Microsoft's Azure status history carries no East US post-incident review for that day, and its most recent entry there is July 23, for West US [15]. Anthropic was willing to name "an issue with an upstream cloud provider" for an August 28 problem affecting Claude Cowork and Claude Code on the web [17], so it attributes when it chooses to. Gemini was the service that kept running while the other three were degraded, and it drew around 500 Downdetector reports without ever posting an outage notice [23] [6]. On the record, that reads as luck rather than any documented difference in architecture.
The exercise worth doing is not a percentage. For each place a model sits in the path of an answer somebody is waiting for, write down the longest continuous outage that workflow can absorb, then hold it against the two durations already on the record: 117 minutes for OpenAI's September 3 incident, 170 for Anthropic's second [7] [8]. Anything that can wait three hours needs a queue and an honest message to the user, not a second vendor. Anything that cannot wait twenty minutes needs a fallback that does not run through the same afternoon, and the option Notebookcheck will vouch for is a model on hardware you control with no internet connection [24]. There is also a cost that rarely makes it into anyone's budget: OpenAI told Codex users on the mobile remote control they may need to pair their device again after the incident [18], which is how a resolved status page becomes another twenty minutes of somebody's Friday.
Ranked by verification strength, evidence, and original report placement.
OpenAI logged an incident titled "Elevated errors across ChatGPT and Codex" at 14:58 UTC on September 3, affecting 15 ChatGPT components and four Codex components, and marked it resolved at 16:55 UTC.
Anthropic recorded an incident affecting Claude Sonnet 5 from 12:37 to 12:56 UTC on September 3.
Anthropic recorded a second incident on September 3 running from 13:26 to 16:16 UTC, covering Mythos and Fable 5.1, Mythos and Fable 5, Opus 5, Opus 4.8 and Opus 4.6.
Google's Gemini was the only one of the four major services without an official outage notice on September 3, although Downdetector still collected around 500 reports for it.
On September 3 ChatGPT peaked at more than 35,000 Downdetector reports in the US, Claude at roughly 1,400 and Grok at roughly 1,200.
OpenAI reports 99.64 percent availability for ChatGPT on status.openai.com over a rolling 90-day window, 99.94 percent for the APIs and 100 percent for Codex.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Scalable Capital puts ChatGPT, Claude and Grok inside the European order ticket2 distinct publishers
product
A school agenda shipped with "Vitoiis" and a planet named Marc, and no one read it first1 distinct publisher
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
product
Half the conversations vanish: an independent check on AI usage reports1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary documents, single reader
Everything quantitative here comes from documents the providers publish themselves — two status pages, two incident histories — and the reporting is unusually careful about what those documents do and do not say. The strongest work is a negative finding: Azure's status history carries no East US entry for September 3, which someone actually went and checked. What holds the score back is that one outlet did all the checking, on a rolling window that has since moved, so the 99.64 and 99.4 percent figures cannot be re-verified from this reporting alone.
Outage traffic is the only demand signal
The nearest thing to usage data is complaint volume: 35,000-plus Downdetector reports for ChatGPT against about 1,400 for Claude and 1,200 for Grok, which tells you the relative size of the crowd that felt the failure and nothing about paying customers, API traffic or enterprise deployments. On the other side of the story, the recommended remedy — a model on your own hardware — has no uptake evidence at all beyond the outlet's own earlier articles.
Cooler than the coverage it corrects
This reporting spends its energy taking a claim away rather than adding one: while other coverage settled on Azure East US, Notebookcheck says the link is undocumented and shows where it looked. That earns a negative reading. Two things pull back toward zero — the 28-times-expected figure treats two providers' aggregate percentages as independent, which is exactly the assumption the piece elsewhere says you cannot verify, and the closing preference for local models is stated as dependability rather than demonstrated.
Self-graded numbers, self-linked remedy
Two incentives sit inside this story and both are visible. The reliability figures are graded by the companies being graded, on pages they maintain, aggregated across tiers and models in a way that flatters any single bad experience away — and the same pages trailed Downdetector on the day they mattered. Then the fix Notebookcheck reaches for routes back through Notebookcheck: its iPhone language-model piece, its Muse Glimmer coverage, its late-August Claude outage report. Neither pressure invalidates the numbers; both explain why nobody outside the vendors publishes a competing figure.
Solid on what happened, empty on why
The timestamps, component counts and percentages can be trusted as faithfully copied; the causal story cannot be trusted at all, because no party has offered one and the reporting says so plainly. That combination — precise facts, missing mechanism, one publisher — is why this lands in the middle. Add that the underlying rolling window keeps sliding, and any figure quoted here has a shelf life.