Science1 distinct publisher2 min readPublished
The odds now favour a record season. The AI forecasting being pitched at the countries with the least margin is the least tested on exactly that kind of event.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Divide the record probability by the very-strong probability and you get the number nobody quotes: conditional on the very strong event arriving, at most about 77% of the probability sits on "bigger than anything measured since 1950" [1]. That comparison window is 76 years long [2]. The central case, in other words, is an event outside the range the instrumental record contains.
This is inconvenient for the forecasting method being pushed hardest at the places with the thinnest margin. AI systems do not integrate the physics; they learn a statistical mapping from the current state of the atmosphere to a later one [5], which makes training expensive and running the finished model cheap enough for ordinary hardware [6]. Models such as GraphCast do beat physics-based systems on speed and general accuracy [7]. General accuracy is an average over ordinary weather. In the tails, the Oxford comment cites studies finding that AI systems underestimate both the intensity and the frequency of record-breaking heat, cold and wind relative to a leading physics-based model [8], and the author notes their reliability has not been properly vetted for extremes under a changing climate [9].
The bias has a map. The author's formulation is that the data gap is most dominant precisely where the vulnerability gap is greatest [11]. Kenya's coastal counties carry the added exposure of storm surge and coastal erosion on top of the rainfall [4].
The recent record also suggests lead time was not the only thing missing. In September 2023 Storm Daniel collapsed two dams at Derna, Libya, killing at least 4,000 people in a single night [12]. In the summer of 2024 the Arba'at Dam in eastern Sudan burst, destroying 20 villages and affecting 50,000 people already living through a civil war [13]. That April the UAE logged its heaviest rainfall in 75 years and Dubai stopped [14]. A sharper five-day forecast helps an evacuation order; it does not raise a spillway. Kenya's own sequence of 1997-98, 2006-07, 2015-16 and 2023-24, each of which killed hundreds [15], puts the last El Nino flood season roughly two years back [3], which is recent enough that the list of unfinished drainage work still exists somewhere.
Two caveats on the material. It is a single expert comment from Oxford, which is itself building an AI forecast product, Forecast4Africa, with AfriClimate AI [16]; the warning about extremes is the author's own, which is more than most product literature concedes. And 69% leaves 31% [4]. An agency that designs only for a record season is designing for the wrong one about three times in ten, which argues for triggered thresholds rather than one planning scenario.
Ranked by verification strength, evidence, and original report placement.
NOAA's Climate Prediction Center has put the chance of a very strong El Nino event this fall and winter above 90%.
NOAA's Climate Prediction Center gives a 69% likelihood that the event will exceed every El Nino since 1950.
In Kenya, high-risk urban centres have already been flagged where poor drainage and strained infrastructure could turn heavy rainfall into a humanitarian emergency.
Kenya's coastal counties face the added threat of storm surges and coastal erosion.
Unlike traditional systems that integrate complex physical equations, AI weather models learn a direct statistical approximation linking current weather conditions to the future.
Training AI weather models requires heavy-duty computation on masses of historical weather data, but once trained they are much cheaper to run and within reach of ordinary computing hardware.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single expert comment relaying uncited third-party figures
All substantive numbers - the NOAA probabilities, the casualty and impact tolls, and the finding that AI underestimates record extremes - come from one republished institutional commentary with no links, dataset references or named studies. The mechanism claims about AI model design and cost are internally coherent and uncontroversial, but the load-bearing evaluative claims (tail-event reliability, data-gap/vulnerability overlap) are asserted rather than demonstrated.
Named projects, no measured operational uptake
There are two concrete programme signals - Forecast4Africa under construction with AfriClimate AI, and the earlier SEWAA work with the World Food Programme across Kenya, Ethiopia, Uganda and Rwanda - but neither comes with users, contracts, forecast volumes, skill numbers or evidence that any national meteorological service currently relies on AI forecasts operationally. Adoption is therefore real but early and unquantified.
Cautionary framing, but promotional and uncited in places
The commentary's substantive direction is deflationary - it argues AI forecasts are least tested precisely on the record-breaking events now expected - which pushes the gap toward zero. It is nudged positive because precise probabilities and casualty figures are relayed without citation, because 'AI may help us better prepare' is asserted on the strength of an unquantified prior project, and because the proposed remedy is the author's own programme. The overstatement is one of confidence and self-positioning rather than of capability.
Author advocates his own institution's programme
The article is an institutional expert comment: the diagnosis (AI models poorly calibrated for African extremes) leads directly to the author's Oxford partnership with AfriClimate AI as the remedy, and to the earlier Oxford/WFP SEWAA work as validation. phys.org's model of republishing institutional communications means no editorial counterweight is applied. There is no funding or conflict disclosure in the supplied text. The incentive is toward programme visibility rather than product sales, which caps the score below the top band.
Plausible and internally consistent, but uncorroborated
The methodological claims about AI weather models and the geography of training data are consistent with how such systems are built, and the derived probability arithmetic is checkable from the source itself. Confidence is held down by one publisher, one author with a stake in the recommendation, no primary NOAA document, and no named studies or benchmarks behind the key reliability claims.
science
Honduras puts 234 of 298 municipalities on drought alert, and names its own end date1 distinct publisher
science
A cheap ligand does the work that heat used to: phenanthroline rewires MOF glasses while molten1 distinct publisher
science
A 3.9C El Nino peak makes 2027 a stress test for plans built for the late 2030s1 distinct publisher
science
What you expect from your own old age shows up a decade later in who you still see1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 26, 2026