Skip to content

Science1 publisher3 min readPublished

ETH Zurich spent a year copying 100 petabytes of NASA data onto servers next to Alps

The Swiss team now has the two things an AI hazard model needs, a full Earth observation archive and a top-tier machine, in the same room. Nobody has reported the forecast skill that pairing is meant to buy.

The Scientist · Science desk

Photograph accompanying ETH Zurich spent a year copying 100 petabytes of NASA data onto servers next to Alps
Photo: swissinfo.ch

What happened

  • ETH Zurich has copied around 100 petabytes of publicly available NASA data onto servers adjacent to Alps, the supercomputer run by the Swiss National Supercomputing Centre in Lugano.
  • Reto Knutti, who heads ETH's Center for Climate Systems Modeling, said the transfer covered roughly 6 billion NASA files and took approximately a year.
  • Knutti put the volume at about 20 million feature-length films, or around a million times the storage on a typical computer.
  • Knutti said a pattern-recognition model can run a multi-day global forecast for the whole planet in about a minute, against the hours a traditional equation-based simulation needs.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost The NASA archive is public, so the bill an agency copying this approach pays is the petabyte-scale store, the year of transfer and the floor space beside a top-tier machine. Those are the items to budget, not the data.
  • capability With the whole record readable at machine speed, searches across it become practical, which is the basis for the claim that these models will surface satellite patterns nobody has been able to see.
  • constraint An early warning system only works where the signal arrives before the event, and Knutti's own framing limits the reliable cases to geological hazards such as landslides and glacier collapses.

Divide 100 petabytes by a year and the copy ran at an average of about 3.2 gigabytes a second, sustained, for twelve months [21]. The file count sets a second pace: roughly 6 billion files over the same period is about 190 files a second, at an average size near 17 megabytes [22] [23]. Neither rate is exotic for a national facility. Both had to be provisioned and paid for before a single model was trained.

Copying at all was a choice. Thomas Schulthess, the ETH computational physics professor who heads the Swiss National Supercomputing Centre in Lugano [2], made the case for locality in terms of wait time. "It matters whether you can move the data within a few seconds or whether you have to wait days for the data to come," he said [8]. The servers holding the archive sit a few meters from Alps [24]. Training models of this kind takes both large volumes of Earth observation data and significant computing power, and ETH says it now has the pair in one place [11].

Reto Knutti, who heads ETH's Center for Climate Systems Modeling, said the models are expensive to train but "cheap to run", so "you can do more iterations, maybe every few minutes, to see if some specific weather pattern exists" [13] [4]. "That allows us to do early warning systems to save lives," he said [14]. The interviews give no figure for the training bill and none for the cost of a single run.

Nature reported last week that satellite image analysis showed some warning signs before the Aug. 26 glacial collapse on the Nepal-China border, and that spotting them could have flagged the area "as a hotspot warranting closer attention and monitoring" [18]. That work was done after the collapse. What it does not say is how many other slopes carried the same signature and stayed intact. That count is what sets a false alarm rate.

Blatten is the closest thing here to a positive control. It was done without AI. Knutti said the Swiss village's destruction by a glacier collapse in May last year "was visible in satellite data more than a year before it actually happened" [16]. The evacuation that avoided mass casualties came a week ahead of the collapse, out of close monitoring by Swiss authorities [17].

Knutti said "Data is essentially everything," and that "The next step will be making sense of the data" [9] [10]. Schulthess called it "a huge scientific opportunity" [6] and said it is "really enabling scientists to do things we would not even have thought of before" [7]. The account given to AFP describes the archive and the machine it sits beside. No forecast produced by these models appears in it.

What to watch

  • A published skill score for an ETH model tested on held-out hazard sites, including the ones where nothing happened.
  • Whether the 100-petabyte copy is kept synchronised with NASA's archive or left to age as a snapshot.
  • Whether other national computing centres budget for their own local mirrors of the same public archive.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories