Science1 distinct publisher3 min readPublished
Google DeepMind says the model resolves temperature at 5 kilometres and updates every hour, and it hands the headline accuracy claim to an outside evaluator, Brightband, whose scores the announcement does not publish.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The training signal is the part worth reading twice. Conventional AI weather models learn to step from one numerical weather prediction analysis to the next, which makes that analysis product their ceiling and hands them its six-hour data lag; DeepMind says the lag shows up as bias in fast-changing variables such as rain and surface temperature [6]. Feeding the network live one-hour geostationary mosaics moves part of the dependence upstream, toward the observations themselves [2]. The satellite feed does not replace the older stream, though: the system diagram keeps traditional historical analysis in the input stack alongside the mosaic [2], so the post's line about learning directly from real-time observations [16] describes a hybrid rather than a departure from reanalysis.
The station work is the other half of the local-scale story. Predicting sparse station coordinates natively [3] is a different task from producing a gridded field, because a gridded field is verified against a smoothed representation of the atmosphere while a thermometer sits in one particular valley. Train on the stations and the 5-kilometre field can carry the topographic offset that coarser models average away [7], which is what the UK temperature comparison in the post is meant to show [15]. It also raises the question the published text does not answer: how station records were divided between training and evaluation [12]. That is precisely why the Brightband attribution carries weight [1]. Scoring forecasts as they are issued is harder to contaminate with choices made in a retrospective rerun, and it is a different exercise from a developer scoring its own hindcasts.
The sharpness arithmetic is worth stating plainly. Going from a 25-kilometre grid to a 5-kilometre grid is a factor of five in each direction, so 25 times as many cells over the same area [13], and moving from six-hourly to hourly cycles means 24 forecasts a day instead of four [14]. What neither number measures is skill at that scale. A generative model can produce a field that looks like terrain-following detail without that detail being verified at 5 kilometres, and the resolution tiers in the announcement quietly concede the limits: wind speed is still a 25-kilometre variable, so this is a 5-kilometre forecast for temperature and moisture specifically [4].
My read is that the resolution figures are the least informative part of the announcement and the satellite-native training path is the claim to track. If a model can be trained largely on raw observations, the barrier for regions the post describes as underserved by high-resolution forecasting [10] stops being supercomputer time for a physics run and becomes access to a satellite mosaic and someone else's inference budget. That is a real change in who can be served, and it will stand or fall on per-variable, per-lead-time numbers that this post leaves to Brightband [12].
Ranked by verification strength, evidence, and original report placement.
The end-to-end WeatherNext 3 system ingests live 1-hour geostationary satellite mosaics alongside traditional historical analysis, feeding a single flexible Functional Generative Network (FGN) mesh transformer.
The model outputs dense gridded fields and discrete cyclone tracks, and predicts station-level sparse coordinates natively.
WeatherNext 3 generates hourly forecasts at multiple resolutions: surface variables such as temperature and moisture at 5 kilometres, other surface variables at 10 kilometres, and atmospheric variables such as wind speed at 25 kilometres.
WeatherNext 2 produced forecasts on a 25-kilometre grid in 6-hour increments, and DeepMind describes WeatherNext 3 as providing a global weather picture roughly five times sharper.
Most AI weather models, including WeatherNext 2, are trained on data from numerical weather prediction models, which are supercomputer-driven physics simulations that carry a six-hour data lag; the post says this lag can lead to biases for fast-changing variables like rain or surface temperature.
WeatherNext 3 trains directly on sparse weather station observation data, which the post says allows global forecasts on a 5-kilometre grid that account for regional details such as topography.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
WeatherNext 3 refreshes a 5 km global forecast every hour off live satellite mosaics2 distinct publishers
build
Anthropic's protein run is checkable, which is rarer than the 26.8% hit rate2 distinct publishers
invest
Meta FAIR says the standard way to plan a training run costs 10x more than it needs to1 distinct publisher
product
WeatherNext 3 moves Google from weather research to weather supplier1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One post, one voice
The specifics are precise and internally consistent — 5, 10 and 25 kilometres, hourly issuance, satellite mosaics plus station observations, cyclone tracks and native station points — but every one of them is Google DeepMind describing Google DeepMind's model. The sentence the whole story hangs on, that this is the most accurate global weather model to date, is credited to Brightband and arrives with no score, no evaluated variables and no method. A UK temperature figure showing crisper topography is a picture, not a measurement.
Launched, not yet used
What exists is a launch, dated the day of the post, plus a promise of reach: reliable forecasts 'accessible across Google products worldwide'. No product is named, no date attached, no external user quoted, no API or dataset pointed at. The clean-energy and underserved-region passages describe who should want this, which is not the same as anyone having it.
Superlative where it cannot be checked
The overstatement and the understatement run in opposite directions. 'Most advanced and accurate global weather model to date' is the strongest claim available in this field, and it is supported by an evaluation nobody can read. Meanwhile 'roughly five times sharper' quietly undersells the step it describes: a 5-kilometre grid holds 25 cells where a 25-kilometre grid held one, and hourly issuance means 24 forecasts a day against four. Google DeepMind reaches furthest exactly where verification is absent, and shortest where the arithmetic is trivial.
The builder is the only witness
This is a product launch published by the lab that built the product, on the day it wanted the news out. The post is pitching consumer reach through Google's own surfaces, and pitching turbine-height wind and solar radiation forecasts to grid operators and renewables developers. Even the independence in the story is second-hand: Brightband is invoked as the outside authority, but it is Google DeepMind that decides what of Brightband's work reaches the reader, and it chose the conclusion without the numbers.
Firm on what, unsure on how good
Two different levels of certainty sit in one story. That WeatherNext 3 exists, ingests live satellite mosaics, trains on station observations and runs hourly at mixed resolutions is stated plainly and consistently enough to rely on. Where it actually ranks against operational forecasts, and whether training on the same sparse stations used to judge it inflates the result, are open questions the announcement does not touch.