Science1 publisher2 min readPublished Updated
NASA recruits volunteers to spot instrument artifacts in real Euclid spectra
NASA's Artifact InSPECtor asks the public to mark cosmic-ray hits, stray glints and electronics quirks in Euclid's spectrograph data, and those classifications will be used to sharpen the software built to remove them.
The Scientist · Science desk

What happened
- NASA has opened a citizen science project called Artifact InSPECtor, which asks the public to help two galaxy surveys, ESA's Euclid and NASA's Nancy Grace Roman Space Telescope.
- The task is to pick out artifacts: signals produced by light glinting off the telescope housing, cosmic rays hitting the detector, or quirks in the camera and electronics rather than by a galaxy or star.
- Volunteers begin on Euclid data, with data from the Roman telescope entering the project in early 2027.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint The tool improves only as far as the human classifications are consistent, so agreement between volunteers looking at the same smudge becomes the practical ceiling on the guidance NASA can extract.
- capability A detector with no earlier archive of its own failure modes gets a supply of human-checked examples soon after its data starts arriving. Those examples are the hardest input to obtain for a first-of-its-kind instrument.
- exposure Distances derived from these spectra feed the dark energy results both missions were built to produce. That puts public classification work upstream of two flagship cosmology measurements.
A spectrograph splits a galaxy's light the way a prism does, and the shape of the resulting spectrum is what gives astronomers the galaxy's distance, the kinds of stars in it, and information about the supermassive black hole at its center [8]. Those spectra are what volunteers will be marking up. A cosmic ray striking the detector, a glint off the telescope's housing, or a quirk in the camera electronics all arrive in the same data as signals that did not come from the sky [4].
Software for this already exists. NASA says astronomers have built AI tools that learn to recognize artifacts, and that recognizing them in data from relatively new instruments is challenging work for the tools, which do not always distinguish them accurately [5]. Euclid, built by ESA with contributions from NASA, is collecting light from millions of distant galaxies now [6]. Roman will capture a similar number after it begins science operations, at different distances and densities across the sky [7].
NASA describes the handoff in a single sentence: the volunteer work "will then be used to improve the instructions guiding the AI tool" [3]. Improving instructions is not the same operation as retraining a classifier on human labels. NASA did not say how accurate the tool is, or how many volunteer classifications each spectrum gets [13].
Volunteers work on Euclid data and will not see Roman data until early 2027 [2]. Whatever guidance comes out of the Euclid classifications will then meet a second instrument with its own detector and its own patch of sky [7].
Euclid's millions of galaxies plus a similar number from Roman put the combined set to be cleaned at several million spectra [12]. Those spectra are how both missions intend to address the expansion of the universe and dark energy [9]. An artifact left in place, or a real feature scrubbed out as an artifact, ends up inside that measurement.
NASA says participants of all ages can take part from a smartphone, tablet, or computer [11]. One of them is nine years old. "It's really cool that we can help teach computers new skills," said Maeve F. after trying out Artifact InSPECtor [10].
What to watch
- Whether NASA publishes how often two volunteers agree on the same spectrum, and how the refined tool scores on spectra no volunteer ever labeled.
- Whether guidance built from Euclid classifications transfers to Roman's detector, or whether a separate labeling campaign is needed for it.
- Whether NASA reports volunteer counts and classification throughput, since that sets how fast the labeling step can keep up with incoming data.