Skip to content

Science1 publisher3 min readPublished

Microsoft Research pipeline forecasts geomagnetic storm risk for 66,935 U.S. substations

Microsoft Research's machine-learning pipeline turns solar-wind forecasts into risk estimates for 66,935 U.S. substations, 30 to 60 minutes ahead. Its scores come from forecasts of fast magnetic-field change, so acting on a warning at a particular substation still needs a field test.

The Scientist · Science desk

Illustration accompanying Microsoft Research pipeline forecasts geomagnetic storm risk for 66,935 U.S. substations
Generated illustration

What happened

  • The pipeline forecasts two geomagnetic indices from solar-wind data measured at the L1 point, then a gradient-boosting model estimates the rate of magnetic-field change at each site.
  • In testing it detected 76.5% of major events, defined as 10 nT/min or more, and 81.2% of severe events at 20 nT/min or more.
  • Its storm-strength forecaster beat the Burton equation on 62.2% of individual hours during the most active periods of 2020-2026.
  • A system of 50 AI agents helped choose features, validation strategies and model settings for the project, built during a Microsoft Research summer internship.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability Tying a storm forecast to latitude, bedrock and conductivity at each substation lets operators single out which sites to protect in the 30 to 60 minutes before impact.
  • constraint With linear regression as the only comparison, the scores show how the model does against a basic fit and leave its value against current utility practice unmeasured.
  • exposure Roughly one major event in four went unflagged, so an operator relying on these site warnings alone would still get no notice of some fast field swings.

The control in this experiment is a straight line. According to the post, no widely deployed operational system makes the same calculation, so there was no industry benchmark, and the risk stage was scored against a simple linear regression [6]. A linear fit on the same inputs tests whether gradient boosting adds anything. It is a modest bar, and the post is plain about why it was chosen.

It also matters what counts as an event. Both thresholds are rates of magnetic-field change: 10 nT/min for a major event and 20 nT/min for a severe one [7]. The post ties that rate, dB/dt, to the risk of geomagnetically induced currents [5]. A fast field swing is still a step short of current flowing through a transformer, and how much current it drives depends on the ground. Regions with resistive bedrock can see stronger induced currents than regions with more conductive geology, and line orientation and latitude shift the exposure of each asset [16].

That geology is why the system works site by site. Its author, a Microsoft Research summer intern [14], wrote: "The challenge is not only knowing that a storm is approaching, but estimating when and where its effects could be most severe with enough warning for grid operators to respond." [15]

The post's summary rounds the major-event detection rate to nearly 80% [8]. The figure behind it is 76.5% [7], so about 23.5% of major events went unflagged [1]. The AE forecaster was designed for the rare, intense activity that drives infrastructure risk [11], and the stronger class was in fact caught more often, 4.7 points above the major-event rate [3].

The Dst model, which describes large-scale storm strength [17], was tested against the Burton equation. Its 62.2% win rate over the most active hours [9] leaves 37.8% of those hours where it did not come out ahead [2]. Its contribution to the full system was small: 1.2 percentage points of severe-event detection [10].

The thing this doesn't tell you is how often a flagged site stays quiet. Every figure above is a hit rate. There is one methodological check to make as well. A system of 50 AI agents helped search features, validation strategies and model settings [12]. With a search that wide, the scores hold only if the evaluation years were kept out of the tuning. Because every input is public [13], outside groups can run that check.

I think the per-site design is the useful contribution. It links a storm forecast to the latitude, geology and ground conductivity under each substation [3]. For now the evidence is an internship project scored on forecasts of magnetic-field change, with a warning window of 30 to 60 minutes [2].

What to watch

  • False-alarm rates published alongside the 76.5% and 81.2% hit rates, showing how often a flagged substation stays quiet.
  • A utility trial comparing the site-level forecasts with currents measured in transformers during a storm.
  • Confirmation that the evaluation years were held out from the 50-agent search over features and validation strategies.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories