Science2 distinct publishers3 min readUpdated
NASA advertises lead times of up to 12 hours. The published test average is 9.24, and the team says the model is not ready for operational forecasting.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
A team working under NASA's COFFIES science center has published a machine-learning model that predicts where solar active regions will emerge up to 12 hours before they become visible on the Sun's surface [1]. That matters because today's operational forecasts, run by NOAA's Space Weather Prediction Center and the U.S. Air Force, start from active regions that are already visible and estimate flare probability from their characteristics [11].
The headline number and the tested number are not the same. NASA describes the capability as up to 12 hours [1]; the study, led by NJIT and published in the Journal of Geophysical Research: Machine Learning and Computation, reports that the best-performing version of the model, called EarlyDetect, identified precursor signals an average of 9.24 hours before active regions became visible in tests on regions it had not seen during training [4][5]. That is a gap of about 2.76 hours between the ceiling and the average [6]. Operators planning a satellite safing procedure or a high-frequency comms workaround should budget against the average, not the ceiling.
The physical premise is indirect observation. Active regions form beneath the photosphere, where the magnetic structure cannot be seen directly, so the model looks for small changes in the magnetic field and in the pattern of acoustic waves travelling through the Sun, which Alexander Kosovichev of NJIT compared to detecting a slight change in rhythm within a very noisy orchestra [10]. EarlyDetect ingests hourly acoustic power maps plus magnetic field measurements from the Solar Dynamics Observatory, with the acoustic maps derived from sound-wave observations that the Helioseismic and Magnetic Imager records every 45 seconds [7]. That is roughly 80 raw samples compressed into each hourly map [8]. The architecture is a sliding-window transformer: instead of looking at all surface activity at once, it moves a fixed-size window across a long timeline, weighting recent data while retaining longer-run patterns [9]. Training and inference used NASA Ames supercomputing resources [22].
The most useful engineering detail is a negative result. The team had applied a filtering step expected to isolate short-timescale patterns, and found it made forecasts worse in almost every case, because it averaged away the faint fluctuations carrying the earliest warning, according to Kosovichev and corresponding author Jonas Tirona, an NJIT undergraduate [14][15][16]. Anyone building a precursor detector on noisy geophysical data should read that as a warning about preprocessing that assumes the signal is the trend.
The operational value depends on how the timing lines up with the physics. Active regions begin emerging over several hours, and full development can take one to several days [17]. Those regions are the main drivers of flares and coronal mass ejections, which send high-energy radiation and charged particles that can threaten astronauts, disable satellites, and disrupt radio communications [24][18]. Tirona said the early warning could let satellite communications or power grid companies prepare and potentially mitigate damage [19]. The shift is from counting visible sunspots to predicting approximate locations of emerging ones [12].
Two caveats are stated by the team itself. The model is not ready for real-time operational forecasting, and the next step is validating it across many more known solar events for tuning [13]. Mengjia Xu, the project's principal investigator at NJIT, said machine learning has not yet been widely applied to solar activity forecasting [20].
Watch whether the 9.24-hour average holds up on a larger validation set, whether NOAA's forecast center shows any interest in ingesting precursor products alongside its visible-region analysis [11][13], and what other groups do with the dataset the team has released [21].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
COFFIES is a NASA DRIVE (Diversify, Realize, Integrate, Venture, Educate) Science Center, and the work brought together researchers from New Jersey Institute of Technology, Princeton University, and NASA's Ames Research Center.
A research team led by NJIT reports in the Journal of Geophysical Research: Machine Learning and Computation that an AI model called EarlyDetect can identify precursor signals of active region emergence in the Sun's acoustic activity and magnetic field.
EarlyDetect analyzes hourly acoustic power maps and magnetic field measurements from NASA's Solar Dynamics Observatory; the acoustic maps are derived from sound-wave observations recorded every 45 seconds by the Helioseismic and Magnetic Imager (HMI) aboard SDO.
The model uses a sliding-window transformer architecture: rather than looking at all activity on the solar surface at once as earlier deep learning approaches did, it moves a fixed-size viewing window across a long timeline of solar activity to focus on recent data while remembering overall patterns.
Alexander Kosovichev, COFFIES co-investigator and distinguished professor of physics at NJIT, said the magnetic structure cannot be directly seen while rising through the solar interior, so the team looks for very small changes in the magnetic field and in the pattern of acoustic waves travelling through the Sun, which he compared to detecting a slight change in rhythm within a very noisy orchestra.
The method allows forecasters to predict approximate locations of emerging sunspots rather than relying on counting already visible sunspots.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed single-study result with held-out test and open data, no independent replication
The core claim rests on a paper in Journal of Geophysical Research: Machine Learning and Computation with a named DOI, a described input pipeline (hourly SDO/HMI acoustic power maps and magnetograms), evaluation on previously unseen active regions, comparison against a standard Transformer and a prior benchmark, and a publicly released dataset and portal that make reproduction possible. Evidence is weakened by being one study from the originating team, described only through the two institutional accounts here, with no independent validation, no reported false-alarm rate, and an acknowledged need to validate across many more events.
Research-stage: data artifacts public, zero operational use
The only concrete uptake facts are artifact releases — the paper and the SolARED dataset plus SAR Portal — and an expression of prospective interest from NASA's Moon to Mars Space Weather Analysis Office. Both publishers state the model is not ready for operational real-time forecasting, current NOAA SWPC and Air Force practice still keys off already-visible regions, and no downstream user, integration, or satellite/grid deployment is reported.
Agency framing overstates lead time and readiness relative to the published test
The measured result is an average 9.24-hour lead on unseen regions from a model the authors say is not deployable, produces false alarms and does not predict flares. The agency framing leads with 'up to 12 hours', says the team 'aims to revolutionize' operational forecasting, ties the work to Artemis and crewed Mars missions, and omits the measured average, the comparison baselines and the false-alarm behaviour. That is a real but bounded overstatement: no source disputes the underlying experiment, and the caveats do appear, just later and unevenly.
Both accounts are institutional self-promotion of publicly funded programs
Both sources originate from the parties that produced the work: one is the funding agency publicising its own DRIVE Science Center and linking the result to flagship Artemis and Mars programs, the other relays a university research communication centred on its own undergraduate author, PI and released dataset. That creates a clear interest in emphasising capability and novelty ('first public dataset', ML 'not widely applied yet') and in framing lead time favourably. Mitigating factors: a peer-reviewed venue, publication of a negative result about the team's own filtering choice, and explicit not-ready caveats. No commercial, vendor or pricing incentive appears in the material.
Two aligned first-party sources, consistent facts, one framing discrepancy
Both accounts describe the same paper and agree on architecture, inputs, institutions, the subsurface detection challenge and the not-ready status, which supports confident assessment of what was done. Confidence is held below high because there are only two sources and both are first-party, the disagreement over the headline lead time is unresolved from the supplied material, and no false-alarm rate, dataset scale or independent evaluation is available.
product
Starling navigated without GPS using the cameras it already had1 distinct publisher
product
A Jupiter-bound probe flew through a CME that Earth-side forecasters had written off1 distinct publisher
product
A five-watt nuclear heater on Blue Ghost turns lunar night into a procurement problem1 distinct publisher
science
GJ 523b gives 'Mega-Earth' a number: 23 Earth masses inside 2.5 Earth radii1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026
1 article · August 14, 2026