Science1 publisher3 min readPublished
Generated tree images help a species classifier only where real photographs are thin
Researchers at NC State mixed AI-generated pictures of 20 North American trees into real iNaturalist and Auto Arborist photos. The synthetic images helped the recognition model only in the runs where real images were limited.
The Scientist · Science desk

What happened
- North Carolina State University researchers tested whether AI-generated images of 20 common North American trees could help train an image-recognition model when real photographs are too few.
- The generated images improved the model's ability to identify trees in the condition the authors set up as scarcity, where real images were limited.
- Taken overall, the synthetic images trained the model less effectively than the real photographs did.
- The work appears in Remote Sensing in Ecology and Conservation, with lead author Thomas Lake of NC State's Center for Geospatial Analytics.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint The account reports no accuracy figures or image counts, so a monitoring team cannot yet read off how many real photographs it needs before the synthetic images stop improving accuracy.
- decision For an under-recorded species with a few dozen photographs, a provisional classifier trained on a real-plus-synthetic mix is now a defensible interim step while collection continues.
- capability The use Lake calls most promising, species and places with little existing imagery, is the one the tested set of common trees could not evaluate.
- exposure Anyone grading a classifier on synthetic-heavy training risks accepting a model that looks right and misses the bark and leaf detail a field identification depends on.
What a generated image can supply is the part a generator has already seen. Lake said "Most of the AI-generated images looked plausible at first glance." Then he added: "But really, they didn't capture the fine details and variations that exist in the real world, telling us that these images can look realistic while still missing subtle information that matters in practice." [7] For trees, the features that separate one species from another include leaf shape, bark texture and growth form. [10]
The shortfall in real data is a matter of coverage as well as counts. "Some species are photographed thousands of times, while others may be rare, occur in remote places or simply receive much less attention. Even for common species, photographs may only come from certain places, seasons or viewpoints," said co-author Chris Jones, a senior research scholar in NC State's Center for Geospatial Analytics. [8] A generator fitted to that same corpus has nothing else to learn from. It cannot add the winter view of a canopy that nobody uploaded.
The test itself ran on 20 common North American trees. Real images came from iNaturalist and the Auto Arborist Dataset, and different mixes of real and synthetic data went into a standard image-recognition model. [1][2] Common species are the easy case for a generator, because they are the ones it will have seen most. Lake put the promise somewhere else: "The greatest potential for this technology may be in studying species and places we know relatively little about," he said. [6] Those species and places sit outside the 20 the study measured.
The phys.org account does not report accuracy figures, the number of real or synthetic images used, or which generative model produced them. [16] A monitoring program that wants to plan around this needs one curve: classification accuracy against the number of real images per class, with the point marked where the synthetic contribution goes to zero. The published title states the condition as real data being scarce, and neither the title nor the release says how scarce. [11]
So the finding is a direction with a known sign and an unknown crossover. That is still useful for a team choosing what to do with 40 photographs of an under-recorded species while fieldwork continues. It is not a reason to spend compute on a class that already has thousands. Lake said the potential should not be overstated. [12] He put the boundary in his own terms: "AI-generated imagery does not eliminate the need for field observations or community science; it gives us new ways to use those observations, for recognizing species and noticing how our Earth is changing, potentially sooner." [9] The work is published in Remote Sensing in Ecology and Conservation. [3]
What to watch
- Whether the published paper plots classification accuracy against the number of real images per class.
- Whether the same low-data benefit holds for animals and insects, where pose and behaviour vary far more than tree form.
- Whether iNaturalist-scale platforms accept synthetic augmentation in the models they run on contributed photos.