Science1 publisher2 min readPublished
Macaque visual neurons follow a motion illusion that current AI vision networks miss
York University researchers found macaque neurons shift their code for a still target's position after motion adaptation, as human perception does. The AI vision networks they tested showed no such shift, so the team proposes the illusion as a NeuroAI benchmark.
The Scientist · Science desk
What happened
- After watching steady motion in one direction, people see a stationary object displaced the opposite way, even though the light reaching the retina has not changed.
- The team paired human perception tests with electrophysiological recordings from macaque inferior temporal cortex, a higher visual area known for object recognition.
- Deep artificial networks often match or surpass humans at classifying objects in static frames, according to the York release.
- Graduate researcher Elizaveta Yakubovskaya led the Current Biology study as first author, with Canada Research Chair Kohitij Kar as senior author.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint Single-image accuracy cannot settle whether a model sees like a primate, because networks that equal people on static frames still miss a shift that IT neurons show.
- decision Researchers who use deep networks as models of IT cortex now have a documented mismatch on moving input, and must add temporal adaptation or limit their match claims to static images.
- precedent As a proposed NeuroAI benchmark, the aftereffect would let groups score models on sequences of frames, so brain-likeness claims built on single images would face a second test.
The illusion is a good probe because it separates the stimulus from the percept. The target stays put and the retinal image does not change [2], so a position code that moves during the aftereffect is following what the viewer sees. Running the same adaptation on people and monkeys [3] means the recorded neurons can be scored against what people say they saw.
The machine side of the comparison needs a closer look. The release describes current deep networks as judging each image through feedforward pixel properties, without the continuous temporal adaptation of animal vision [6]. A network that sees the target frame with no record of the frames before it has nothing for the adapting motion to act on. On that description, its failure to shift [5] is close to what the design predicts. I think the result is still worth having: it turns a known gap in the architecture into a behavioural test with a pass or fail answer.
The press account does not say how many monkeys or neurons were recorded, how large the shift was, or which networks were tested. Those details decide whether the release's "current state-of-the-art AI vision networks" [5] includes video or recurrent models that carry information from frame to frame, or only single-image classifiers.
The benchmark measures brain-likeness, and that is a separate target from accuracy. The York release describes perceptual errors of this kind as signatures of optimal, energy-efficient biological computation [12]. It argues that AI meant to work safely and intuitively alongside people must be trained on these history-dependent computations as well as on static pixel accuracy [11]. The study as described documents a gap between brains and models and does not test whether closing it improves any task. In the aftereffect the object has not moved [2], so a pixel-bound network reports its physical position correctly. For a model built to predict what a person will perceive, the shift is what it should produce. For a model built to find objects, the unshifted answer is the right one.
"Today's AI vision systems are impressive, but they still do not always see the world the way we do," said Kohitij Kar, the study's senior author [8] [9]. "By using smart experiments to reveal the computations biological vision uses and AI still lacks, we can use those insights to build better, more brain-like artificial systems," he said [10].
What to watch
- Whether the Current Biology paper's network comparison includes video or recurrent models that carry information across frames, and whether any of them shift.
- The number of monkeys and neurons recorded and the size of the IT position shift, as reported in the full paper.
- Whether other NeuroAI groups adopt the aftereffect as a benchmark and publish scores for sequence-processing models.