AI Summaries Quietly Distort What Readers Remember

Misleading AI summaries cut recall of a key detail from 83.6% to 44.8% in a Georgetown study. Labeling the source as AI or human did not change the effect.

Read a misleading AI summary once, and your memory of the original can shift - sometimes by almost forty percentage points - even when you know a chatbot wrote it. That is the punch of a controlled experiment out of Georgetown, accepted for the AAAI/ACM Conference on AI, Ethics, and Society (AIES) this October and reported by Georgetown on September 25. It lands at a moment when generative AI is being pushed into exactly the workflows where memory matters most: police body camera review, medical records, courtroom filings, internal investigations, newsroom triage.

What the Study Found

The team - Mattea Sim of Georgetown’s Massive Data Institute, Yoshi Kohno of Georgetown’s Computer Science department, and Yael Eiger, a PhD candidate at the University of Washington - ran a two-part experiment with 331 participants who completed both sessions. In the first part, they fed animated traffic-stop videos to OpenAI’s ChatGPT and Google’s Gemini and asked the models to summarize them. In the second, they showed those same videos to human participants, then handed them a summary that was either accurate or had been deliberately altered, and asked them 24 to 48 hours later to remember a core detail: was the sign a stop sign or a yield sign?

The headline result is not subtle. Among participants who read the accurate summary, 83.6% answered the sign question correctly. Among participants who read the misleading summary, only 44.8% got it right - a drop of nearly forty percentage points for a single, central fact. Telling readers that the summary had been written by a chatbot rather than by a human transcriber made no measurable difference. Telling them the summary was human when it was AI also made no measurable difference. The contamination was indifferent to provenance.

The summaries themselves were bad in ways that matter. On average, AI-generated summaries omitted 51.6% of the central details in the underlying video. All but one of the tested summaries failed to mention the most central event - a car colliding with a pedestrian. Yael Eiger, one of the co-authors, told Georgetown that she was “struck by how bad the summaries were, even at this stage in AI development.” Mattea Sim put the policy point bluntly: “AI is a new method of delivering misinformation, and it has the potential to create these false memories for people who are reading that information.”

Why “Humans in the Loop” Is Not the Fix

The classic defense of AI summarization in high-stakes settings is that a human will catch the errors. The Georgetown paper turns that defense inside out. “Though ‘humans-in-the-loop’ are often expected to correct for AI’s mistakes,” the authors write in the preprint on arXiv, “our work suggests human memory can instead be distorted by these mistakes.” A reviewer who has just read a confident AI summary does not bring a fresh recollection of the original to the task; they bring an AI-shaped recollection, and they read the source through it.

The study’s follow-up plan is to extend the same experimental setup from animated traffic clips to real police body-worn camera footage. Sim told Georgetown she wants to know: “With police body camera footage, what happens with the AI summaries and what do those police reports look like when they’re generated by AI? Does this manipulate people’s memories of police-civilian interactions and perhaps consequently their perceptions of what happened in the incident?” That question is not hypothetical. Body-worn cameras were introduced about 15 years ago, police departments discard most footage after roughly 90 days unless tied to investigations, and a separate Princeton Engineering study tested 12 general-purpose AI models - including Gemini 2.5 Flash, GPT 4.1, Llama-VID, and Qwen 2.5VL - against 185 hours of public bodycam from Illinois, California, Texas, and Washington, D.C., and found that the models failed to identify basic action details in one out of four cases. The most accurate model hit roughly 77% on one-minute clips; the least accurate hit 11%.

The Privacy Hook

For a site that tracks AI privacy, the implication is concrete. The Georgetown result is the first controlled human-subjects evidence we have that an AI-generated summary is not just wrong but actively rewrites the recall of a reader who never asked the AI for an opinion. That is a different failure mode than a hallucinated statistic in a research report or a wrong answer in a code review. It is a memory-editing failure. The reader leaves the experience holding a different past than they walked in with, and they usually do not know it.

The manipulation surface goes the other direction, too. A Cipher Brief piece from August by former CIA analyst Candice Bryant and former Senior Foreign Service officer John Pennell documents “Generative Engine Optimization” (GEO): the deliberate shaping of web content so chatbots absorb and repeat it. The authors cite a NewsGuard finding that leading AI models repeated narratives from a Russian-linked network roughly a third of the time, and a Harvard Misinformation Review study putting the number closer to 5%, clustered in “data voids” - obscure topics where credible sources are thin. Either number is enough to corrupt the summary that an analyst, journalist, or juror relies on. Add the Georgetown result, and a reader who goes to the AI for help parsing a messy record can end up with a memory of the record that has been filtered, condensed, and quietly tilted before they ever form an opinion.

What This Means

The findings sit in a tight pattern. AI summarization is being adopted fastest in the places where the input is too large for a human to read in full: body camera review, surveillance logs, medical records, legal discovery, newsroom triage. Those are also the places where a confident wrong summary does the most damage, because there is no time to second-guess it and no appetite to. The Princeton team’s conclusion - that general-purpose tools like ChatGPT and Gemini are not ready for this - aligns with the Georgetown result. Both point at the same gap: high-stakes summarization is being deployed by agencies that have neither the time nor the audit rights to check what the model is leaving out.

For users, the practical move is to read the source before reading the summary, or to read them side by side. For the people shipping these systems, the move is harder: stop selling AI summaries of high-stakes material as a productivity feature, run them as a structured workflow with a human who has seen the original first, and treat omission rates - not just accuracy on multiple-choice benchmarks - as a release-blocking metric.

The Bottom Line

A 331-person Georgetown experiment showed that misleading AI summaries drop a reader’s recall of a key video detail from 83.6% to 44.8%, and that labeling the summary as AI-generated offered no protection. Combined with separate evidence that general-purpose AI models already miss basic details in one of four bodycam clips, the privacy and evidentiary risk of AI summarization in police, medical, and legal workflows is real and current, not hypothetical.