Science1 distinct publisher3 min readUpdated
A Physical Review Letters paper reports that AI plus rare event sampling matched a 50,000-run heat wave study using roughly one-hundredth as many simulations.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
An international team in the United States and France has published a hybrid forecasting method, AI+RES, in Physical Review Letters, aimed at the one thing machine-learning weather models reliably get wrong: events that were not in their training data [1][8]. The work was co-led by members of the Climate Extremes Theory and Data Group at the University of Chicago, run by associate professor of geophysical sciences Pedram Hassanzadeh [1].
The trade-off being attacked is familiar to anyone who has costed out a compute budget. Physics-based supercomputer models can produce genuine extremes, but need large amounts of time and energy [2]. AI models are fast and good at day-to-day forecasting, and then fall over on outliers [3]. Hassanzadeh puts it bluntly: AI weather and climate models are "one of the great achievements of AI in science, but they're not magical - they fail on gray swans, the rarest and most extreme events," while detailed physics-based models capture extremes at prohibitive cost [4].
The cost problem is a sampling problem. To estimate the odds that Chicago hits 90F (32C) in July you do not need many simulations, because it is not uncommon; to estimate the odds of 105F (41C) you need far more runs before the extreme shows up at all [5]. Rare event sampling (RES) reduces that bill by scoring conditions so the climate model spends its cycles only on the promising ones [6]. Its known weakness is duration: RES does not work well for short events such as a weeklong heat wave, as opposed to a whole season that runs unusually hot [7]. AI+RES replaces the scoring step with an AI prediction of which conditions lead to shorter, rapidly developing extremes [8]. Run iteratively, according to co-first author Alexander Wikner, a Schmidt AI in Science Postdoctoral Fellow in Hassanzadeh's group, you end up with a set of simulations that do capture the rare event, and the more of them you have the better the probability estimate [9].
The benchmark is the part worth noting. The team ran 50,000 simulations with a traditional climate model to predict heat waves over parts of France and the U.S. Midwest, and reports that AI+RES produced nearly identical results with one-hundredth as many simulations [10] - on the order of 500 runs [11].
The stakes are not abstract. The 2003 European heat wave was associated with roughly 70,000 deaths, and Russia's 2010 event with 56,000 [12]. This past June, nearly half of the United States, about 180 million people, saw dangerous temperatures [13].
Two caveats sit inside the paper's own framing. This was a proof of concept run on a model that does not account for climate change, and the researchers say they hope to test it under different warming scenarios [14]. And Wikner notes the method could be used to generate rare-event data sets to train better AI models, which would in turn speed up the method [15] - a loop that is attractive and also circular, since the AI doing the scoring would then be trained on synthetic extremes.
What to watch: whether the hundredfold reduction survives a model with climate forcing in it, whether AI+RES holds up on extremes other than heat waves, and whether anyone reproduces the France and Midwest result independently. Until then this is one benchmark on one variable, published by the group that designed the method.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
An international team of researchers in the United States and France, co-led by members of the Climate Extremes Theory and Data Group of Pedram Hassanzadeh, University of Chicago associate professor of geophysical sciences, developed a new hybrid method published in Physical Review Letters.
Traditional supercomputer-based physics models can forecast once-in-1,000-year events but require a lot of time and energy.
Newer AI-based forecasting models are good at day-to-day forecasts but often fail to predict outlier events that were not represented in their training data.
Hassanzadeh said: "AI weather and climate models are one of the great achievements of AI in science, but they're not magical - they fail on gray swans, the rarest and most extreme events. Detailed physics-based models can capture extremes, but they require prohibitively large amounts of time and energy."
Estimating the odds that Chicago reaches 90F (32C) in July, which is not uncommon, requires few simulations; estimating the odds of 105F (41C) requires many more runs before that extreme appears.
Rare event sampling (RES) is a statistical technique that speeds up simulation by scoring conditions so the climate model focuses only on the most promising and ignores the rest.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed result, single-source relay
The core claim rests on a named, peer-reviewed Physical Review Letters paper with a DOI and a quantified comparison (50,000 runs versus one-hundredth as many), which is stronger than an unpublished announcement. It is weakened by coming to Clarity through a single institutional writeup with no methodological detail on the AI component, no accuracy metrics beyond 'nearly identical results', and no independent replication.
Research-stage, no deployment reported
Adoption evidence extends to a journal publication and one in-house benchmark on heat waves over parts of France and the U.S. Midwest. The source explicitly frames the work as a proof of concept using a model without climate-change forcing, and describes connecting a state-of-the-art numerical prediction model to a trained AI model as the next step, so no operational or third-party use exists in the supplied material.
Modestly overstated relative to stage
The reported 100x simulation reduction and the framing that the framework is 'ready to be scaled' for regional decision-making run ahead of a single author-run benchmark on an idealized model without climate-change forcing, with accuracy described only as 'nearly identical'. The gap is moderate rather than severe because the underlying result is peer-reviewed and the source itself carries the proof-of-concept caveat and the AI-fails-on-gray-swans limitation.
Institutional research promotion
The only source is an institutional research writeup relayed by phys.org, quoting the principal investigator and a fellow from his own group about the significance and scalability of their own method; the named Schmidt AI in Science fellowship affiliation also signals a funded research-visibility interest. No commercial vendor or product incentive is evidenced, which keeps this below the top of the range.
Single publisher, one benchmark
Confidence is limited by one publisher, one source item, and one author-run benchmark, with no independent verification of the efficiency ratio or accuracy parity. It is not lower because the underlying artifact is a peer-reviewed paper with a DOI, the numbers are specific, and the source states its own scope limits.
science
A centuries-old Coulomb's law test could out-search accelerators for millicharged particles1 distinct publisher
science
A sensor tuned to its own noise beats the textbook entangled state by 0.698 dB1 distinct publisher
science
Tungsten's damage curve has a bump in it, and fusion lifetime models miss it1 distinct publisher
invest
Pleasant, and lonelier: a 12,365-person trial cuts against the AI companion pitch1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026