Science1 distinct publisher3 min readPublished
A five-season Atlantic study led from the University of Miami found ECMWF's AI ensemble carrying more tropical cyclogenesis signal at 84 to 120 hours and less at 36 to 48, which makes it guidance rather than a substitute.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Eighteen developing systems in a single season is a thin denominator for any statement about lead time, and it is the denominator holding up the per-lead-time comparison [5]. The reversal is the part worth keeping: the direction of the AI-minus-physics difference flips as the forecast shortens, positive at roughly three and a half to five days and negative inside two [19][15][16].
The thing this does not tell you is the false alarm side. Probabilities reported for waves that went on to become storms measure only the hit side, and a system that is liberal with genesis probability will look good on that set by construction [5]. What would settle it is how often each ensemble raised probabilities on African easterly waves that never organized [3]. Reliability, not generosity, is what a forecaster is buying at day four.
The record is not stationary either. ECMWF's conventional ensemble improved over the period, its average genesis probabilities generally rose across the five seasons examined, and its grid spacing was halved partway through, in 2023 [4][18]. Halving spacing in both horizontal directions means roughly four times as many grid columns over the same area [17], so the multi-year trend mixes a changing model with changing weather and should not be read as a clean skill curve. The timing does help in one respect: the 2024 comparison sits after that upgrade, so the AI ensemble was measured against the sharper physics ensemble [4][5].
On position, the AI ensemble mean was often closer than the AI single forecast, the deterministic IFS and the IFS ensemble mean [9]. Ensemble means are averages and averaging suppresses error, so the entry in that list that carries real weight is ensemble mean against ensemble mean.
Genesis is the right place to look for this because the signal is unstable by nature. The study's own framing is that the indication a disturbance will become a cyclone can change substantially as the event approaches [14], and the conventional probabilities often jumped sharply between three days out and two days out [7]. A late jump is information arriving when there is least time to use it. Extra signal at day four on the stronger waves buys monitoring time rather than accuracy, which is what Majumdar claims for it: better identification several days ahead "could give forecasters more time to monitor disturbances, assess possible tracks and communicate emerging risks" [12].
Where I will commit, with the conditions attached: on stronger African easterly waves at three and a half to five days, in one Atlantic season, the AI ensemble carried genesis signal the physics ensemble did not, and how much varied storm to storm depending on the wind-speed threshold used to define development [5][8]. Majumdar's own reading is that this is complementary guidance rather than a replacement for numerical prediction [10], and the short-range behaviour on weaker systems is the reason to take that literally [6].
Ranked by verification strength, evidence, and original report placement.
A study led by Sharanya J. Majumdar, professor of atmospheric sciences at the University of Miami Rosenstiel School of Marine, Atmospheric, and Earth Science, compared ECMWF's Integrated Forecasting System (IFS) with its Artificial Intelligence Forecasting System (AIFS), published in Weather and Forecasting (2026), DOI 10.1175/waf-d-25-0200.1.
The researchers analyzed African easterly waves and tropical cyclone development across the Atlantic from 2020 through 2024.
African easterly waves are westward-moving atmospheric disturbances that form over sub-Saharan Africa during the boreal summer, and some eventually organize into tropical cyclones.
ECMWF's conventional forecasting system improved over the period examined: in 2023 the grid spacing of the IFS ensemble was reduced from 18 kilometers to 9 kilometers, while average probabilities of tropical cyclone formation generally increased over the years.
Among 18 tropical cyclones that developed in 2024, the AIFS ensemble frequently produced higher probabilities of development than the IFS at lead times of roughly 84 to 120 hours, particularly for stronger tropical waves.
At shorter lead times of 36 to 48 hours, the AIFS ensemble generally produced lower probabilities of development than the IFS, especially for weaker systems.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Indonesia's peat fires already match big El Nino years, and the forecast peak is October1 distinct publisher
science
NOAA's 69% record El Nino call lands where forecast models are weakest1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed, multi-season, but no magnitudes and one publisher
The underlying work is a peer-reviewed paper in Weather and Forecasting with a DOI, spanning five Atlantic seasons and co-authored with ECMWF, NCAR and University of Bonn scientists, and it reports directional results across two lead-time windows plus position-error comparisons over four model configurations. What holds the score down: the probability comparison rests on 18 storms in a single season, the source reports no probability values, error magnitudes, skill scores or uncertainty bounds, results are acknowledged to swing with the wind-speed development threshold and from storm to storm, and only one publisher account is available for verification.
Both systems operational at ECMWF; desk uptake unmeasured
Adoption is real at the model-production layer: AIFS has been operational at ECMWF since 2025 and the physics baseline received an operational resolution upgrade in 2023, so the compared systems are production systems rather than research prototypes. There is no evidence in the supplied source of downstream uptake - no forecaster or national-center usage disclosure, no case of AIFS genesis guidance changing an advisory or warning, and no user counts - so the score reflects producer-side deployment only.
Framing tracks the evidence, marginally understated
The account is unusually disciplined for AI-in-science coverage: the headline offers 'guidance', the lead author explicitly rejects replacement of numerical weather prediction, the short-lead disadvantage of the AI system is reported alongside its long-lead advantage, and improvements in the physics baseline are credited. The only inflation pressure is qualitative language about 'substantial' potential significance unaccompanied by numbers. Because a directional, peer-reviewed operational result at a production forecast center is presented in deliberately hedged terms, the framing sits at or slightly below what the evidence supports.
Institutional promotion plus self-evaluation by system owners
The single account is research-communications material centered on a named professor at his own institution, a genre that selects for favorable framing of that institution's work. The evaluated systems' operator, ECMWF, is a co-author on the study, so the assessment is partly self-assessment of both the AI and physics products. No independent forecasters or competing centers are quoted. Mitigating factors: peer review, and the reporting of a finding unfavorable to the AI system at short lead times, which cuts against pure promotion.
Credible and specific, but single-publisher and unquantified
Confidence is moderate: the specifics are checkable (DOI, journal, named institutions, explicit lead-time windows, dated operational milestones) and the framing is internally consistent, which supports the qualitative picture. It is held back by having exactly one publisher in the cluster, by the absence of any numeric magnitudes to audit, by the small 18-storm probability sample, and by author-acknowledged threshold sensitivity and inter-storm variance.