Science1 publisher3 min readPublished
Misleading summaries rewrite what eyewitnesses recall of a crash even when readers know AI wrote them
Georgetown and University of Washington researchers found misleading AI summaries cut recall of a crash video's key details from 83.6% to 44.8%. Knowing the text came from AI did not protect readers who had watched the event themselves.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The team first had ChatGPT and Gemini summarize animated traffic-incident videos, and the models left out an average of 51.6% of central events.
- In 95% of tested runs, the summaries omitted the most consequential event in the footage, a car hitting a crossing pedestrian.
- For the memory test, 331 participants watched a red car at a stop or yield sign turn and hit a pedestrian, then read an accurate or altered summary 24 to 48 hours later.
- Participants absorbed the errors regardless of how familiar with or trusting of AI they said they were.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint Disclosure labels cannot be counted on to protect memory: in this experiment, telling readers a summary was machine-written left the recall loss in place.
- exposure Where a digest of camera footage or case notes becomes the record, its errors can pass into the testimony of people who saw the event, where they can look like independent corroboration.
- decision Agencies trialling video summarization have to check whether a tool captures the central event at all before anyone reads its output, since consumer models missed the collision in 95% of runs.
The memory experiment crossed two variables: whether the summary was accurate, and who readers were told had written it, an AI model or a human transcriber [7][8]. Everyone watched the same short event, a red car at a stop or yield sign turning into a pedestrian [6]. That makes the accurate-summary group the control. After one to two days and a faithful digest, those participants got key scene details right 83.6% of the time [9]. Readers of an inaccurate AI-labeled summary got 44.8% right [9]. The gap is 38.8 percentage points [1], a relative drop of about 46% [2].
The provenance arm tests disclosure as a fix. If people discount text they know a machine wrote, the AI label should have shrunk the effect. According to the researchers, it did not, and participants took on the errors whatever their stated familiarity with or trust in AI [10]. The work was presented at the Ninth AAAI/ACM Conference on AI, Ethics, and Society [1]. The Neuroscience News account does not report recall for misleading summaries labeled as human-written, or how the 331 participants were divided among the conditions [6].
The study's two parts tested different kinds of error. The audit of ChatGPT and Gemini mostly found things missing: 51.6% of central events on average, and the collision itself in 95% of runs [3][4]. The memory test used summaries with altered, erroneous details [7]. So the evidence shows that wrong details in a summary displace what people remember seeing. Whether a summary that simply drops the collision has the same effect is a separate question. Invented details are part of what the real tools emit, too: the models also produced recurrent hallucinations [5].
"I was struck by how bad the summaries were, even at this stage in AI development," said co-author Yael Eiger, a Ph.D. candidate at the University of Washington [13]. Eiger also said: "It worries me that police departments may be using video summarization technologies without rigorous testing and without an awareness of how incorrect AI-generated summaries could be" [14]. Lead author Mattea Sim, an assistant research professor at Georgetown's Massive Data Institute, said: "AI is a new method of delivering misinformation, and it has the potential to create these false memories for people who are reading that information" [11].
The study leaves open how well the result carries over to the settings Eiger names. The clips were animated versions of classic eyewitness paradigms, the summarizers were consumer chatbots [2], and the delay was fixed at 24 to 48 hours [7]. Institutions are putting language models on body-worn camera logs and clinical case notes [15]. A test of general-purpose models on animated crashes is a research benchmark, and those deployments run different products on messier footage.
On human review, I think the study supports a narrower claim than the idea that a reviewer is no safeguard. Its participants saw an event and read a digest of it a day or two later [6][7]. An officer who attended an incident and later reads an AI digest of the camera footage is in that position. A reviewer who checks the summary against the footage before signing off is a different case, and the design described here did not test it [7].
What to watch
- The full AIES paper's recall figure for misleading summaries labeled as human-written, and how the 331 participants were split across conditions.
- A replication using real body-worn camera footage and the summarization products police departments use, instead of consumer chatbots on animated clips.
- A test of whether a summary that simply omits the collision shifts recall the way planted wrong details did.