Science1 publisher2 min readPublished
Deep networks grouped toddler cries by sound and missed which message each carried
A Tel Aviv University team gave one classical acoustic method and two deep neural networks a set of toddler recordings whose meanings adult listeners could confirm. All three grouped the clips by sound.
The Scientist · Science desk

What happened
- A Current Biology study led from Tel Aviv University argues that AI models capture the physical properties of a sound without capturing the meaning the receiving animal assigns to it.
- To get a case with a checkable answer, the team used vocalizations of toddlers who have not yet fully developed speech, because the adults those sounds are aimed at can report how they interpreted them.
- The clips were sorted by a classical acoustic method and by two deep neural networks, one trained on animal vocalizations and one on adult human speech, each asked to group the sounds by their characteristics.
- The models put together vocalizations that carried different messages, and split apart different vocalizations meant to convey the same message.
- They also failed to identify how a sequence of vocalizations expressed increasing urgency, which the study says the human ear perceives naturally.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint A vocabulary assembled only from acoustic clustering has no internal check on itself. The categories it publishes are categories of sound, and no step in the pipeline tests whether the receiving animal treats them as different messages.
- cost The validation the authors call for is per-species animal work: behavioral observation, playback trials and sometimes neural recording. A larger model does not reduce that cost, and it does not transfer from one species to the next.
- capability The toddler set-up gives the field a scoreable test bench. A new model can be measured against known listener interpretation before anyone points it at a whale population where no such answer exists.
Animal-call research runs without an answer key. Attempts to decode bats, whales and birds with AI have multiplied in recent years [21]. A recording on its own does not say whether the animal hearing a call treats it as the same request as the last one. The toddler clips came from three situations: distress, calling for the mother or father, and asking for food [7].
Both deep networks did better than the classical acoustic analysis. Neither recovered the messages [9]. A better acoustic representation improves resolution on the property being measured. The researchers say the loss is on the receiver's side: acoustically similar sounds need not carry similar meanings, and sounds that do not resemble each other can carry the same information to whoever hears them [4].
The two error directions do different damage to a published vocabulary. One inflates the count of call types; the other hides a distinction the receiver is making [10]. The material here was human. The claim that acoustic clustering "may create a misleading picture of the communication system and the meaning of the messages it conveys" is an argument the authors build from the one case where the picture can be checked [5]. The phys.org account did not disclose how many toddlers or clips were involved, or how well any of the three pipelines [19] recovered meaning.
The framing the authors offer is perceptual: every species has its own perceptual world, so understanding what an animal is saying means examining how it hears the sound and how it responds [16]. The paper, "The challenge of decoding animal communication using AI", appeared in Current Biology [17]. It has six authors across five institutions [18]: Tel Aviv University, the Hebrew University of Jerusalem, the University of Edinburgh, the Museum für Naturkunde-Leibniz Institute for Evolution and Biodiversity Science, and Humboldt-Universität zu Berlin [3].
Yovel said, "In recent years, there has been growing excitement about the possibility of using artificial intelligence to decode animal communication, but our study shows that these promises should be treated with caution" [13]. He added: "Artificial intelligence is a powerful tool, but it is no substitute for the perspective of the animal itself" [14].
What to watch
- Whether the full Current Biology paper reports sample sizes and cluster-agreement scores that show how large the failure was.
- Whether any group re-runs a published animal call catalogue through a ground-truth check of this kind, or publishes playback validation alongside its clustering.
- Whether grant reviewers start asking animal-decoding projects for receiver-side evidence, meaning playback trials or neural recording, before accepting a claimed vocabulary.