When AI Summaries Mislead: Eyewitness Memory at Risk

Summary:

Generative AI summaries can distort eyewitness memory, leading people to misremember events they personally observed. In a controlled experiment, exposure to misleading AI-generated summaries substantially reduced participants’ recall accuracy and implanted false details—even when readers were told the summaries were produced by AI.

Key Facts:

  • High Omission Rates: Across commercial multimodal models, including ChatGPT and Gemini, automated summaries omitted an average of 51.6% of central events. In 95% of cases tested, the summaries failed to mention the most critical detail: a car colliding with a pedestrian.
  • Recall Accuracy Dropped Sharply: Participants who read accurate summaries recalled key scene details correctly 83.6% of the time, while those who read inaccurate AI summaries achieved only 44.8% accuracy.
  • “Human-in-the-Loop” Is Not a Guaranteed Safeguard: Labeling a summary as AI-generated did not prevent memory contamination; participants accepted erroneous details regardless of whether they believed the text came from a human or a machine.

Source: Georgetown University / University of Washington

Overview: Institutions increasingly use large language models to summarize long video and audio streams—from corporate meetings and clinical notes to law enforcement body-camera footage. But decades of cognitive research show that human episodic memory is reconstructive and vulnerable to post-event misinformation. A joint study by Georgetown University and the University of Washington examined whether AI-generated summaries can introduce such distortions.

Presented at the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES), the study found that misleading AI summaries systematically altered eyewitness recollection, even when readers knew the summaries were produced by an automated system. Lead author Mattea Sim, assistant research professor at Georgetown’s Massive Data Institute, warned that AI is a novel method for distributing misinformation that can create false memories for readers.

Missing the Collision: Widespread Omission in AI Summaries

The researchers evaluated two key aspects of AI integration: the factual fidelity of commercial models when summarizing video, and the cognitive impact of those summaries on human observers. They prompted models—including OpenAI’s ChatGPT and Google’s Gemini—to summarize animated traffic-incident videos based on classic eyewitness-reliability paradigms. The outputs showed frequent hallucinations, notable descriptive gaps, and systematic omissions.

On average, models omitted 51.6% of the scene’s central events. Most strikingly, in 95% of iterations the summaries left out the most consequential element: a motor vehicle striking a pedestrian. Yael Eiger, a study co-author and Ph.D. candidate at the University of Washington, expressed concern that law enforcement agencies may be adopting summarization tools without fully understanding their error patterns or the downstream risks.

How Misleading Summaries Alter Memory

To measure cognitive effects, the team recruited 331 participants who watched animated videos showing a red car approaching an intersection controlled by either a stop sign or a yield sign, then turning and hitting a pedestrian. After 24 to 48 hours, participants read a narrative summary of the incident: some received accurate summaries, others read versions that contained altered or incorrect details. The experiment also varied participants’ beliefs about the summary’s origin, telling some that the text was generated by AI and others that it was produced by a human transcriber.

When memory was tested, results showed a stark difference between groups. Participants who read accurate summaries scored 83.6% on recall assessments, while those exposed to misleading summaries scored only 44.8%. Importantly, warning readers that the summary was AI-generated did not reduce the effect: belief about provenance, baseline trust in AI, and prior familiarity with AI technologies did not protect participants from accepting false details or experiencing degraded memory performance.

Implications for Policy and Practice

The findings challenge the assumption that human oversight is a sufficient safeguard against AI errors. While human review is widely promoted as a corrective layer, the research suggests that humans who read mistaken summaries may internalize those mistakes and carry them into subsequent reports or testimony. In policing, judicial, and other high-stakes settings, an officer or eyewitness who reviews an erroneous AI summary before writing a statement could unintentionally adopt those inaccuracies as genuine memories.

The authors emphasize that AI can generate misinformation even without malicious intent, and that such misinformation can meaningfully affect human memory. Yoshi Kohno, McDevitt Chair in Computer Science, Ethics, and Society at Georgetown and a co-author, described the study as part of broader work examining the interaction between AI systems and human cognition, grounded in psychology and computer science.

Next steps for the research team include moving from synthetic animations to real-world body-worn camera footage to assess how commercial summarization tools might alter official reports and civilian testimony in actual legal contexts.

Editorial Notes:

  • This article was edited by a Neuroscience News editor.
  • The journal paper was reviewed in full.
  • Additional context was added by staff.

About this AI and memory research:

  • Media Contact: Jason Shevrin
  • Source: Georgetown University
  • Image Credit: Image credited to Neuroscience News
  • Original Research (Open Access): arXiv (Sept 23, 2026). Title: “AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory.” Authors: Mattea Sim, Yael Eiger, and Tadayoshi Kohno.
  • DOI: 10.48550/arXiv.2609.28820

Abstract

This study assessed whether errors in AI-generated video summaries distort human memory. First, the authors analyzed summary outputs from large language models to quantify common error types, finding frequent omissions of critical details. Second, a human-subjects experiment tested the downstream effects: participants watched a car-pedestrian incident and later read either an accurate or a misleading summary. Those who read misleading summaries were significantly less likely to recall the original event accurately. The results highlight risks when AI summarization is used in high-stakes domains and call into question reliance on human review alone as a safety mechanism.