Sixty-seven people signed up to spend four weeks judging whether news headlines and images were real. By the end, the ones who had leaned hardest on a chatbot to make those judgments were the ones doing worst without one. That is the headline finding from a controlled study out of the MIT Media Lab, peer-reviewed and published in April at the CHI ‘26 conference on Human Factors in Computing Systems, and re-surfaced this week in MIT Technology Review’s feature on “Your brain on AI”.
It is a small but careful experiment, and it lands at a moment when a quarter of young US adults have used a chatbot for news at least once, according to Pew data cited in the MIT News write-up. If the result generalizes, the fastest-growing way to “check” a story is quietly eroding the underlying skill.
The Setup
The team, co-led by MIT Media Lab PhD students Anku Rani and Valdemar Danry with senior author Pattie Maes, recruited 67 participants in the United States and United Kingdom and ran them through roughly 50 validated news items, mixing text headlines with paired images. The task: is this real, or is this fake? Over four weeks, each participant completed repeated sessions, sometimes with a chatbot assistant (ChatGPT or Claude) on the screen, sometimes without.
The setup lets the team measure two distinct things: how much the chatbot helps in the moment, and what happens to the human’s judgment when the chatbot goes away.
What the Numbers Show
With the chatbot on, participants were 21% more accurate at flagging fake news than they were working alone. That is the result that chatbot vendors tend to lead with: faster, better fact-checking, available on demand, in your chat window.
The number that does not show up in vendor decks is what came next. By the end of the four-week study, the same participants were 15 percentage points worse at spotting fake news on their own than they had been at baseline. They had used the tool, gotten better with it, and gotten worse without it. About a quarter of participants self-reported feeling more confident in their own judgment by week four, even as their measured accuracy had fallen. The team calls this the “AI dependency paradox”: a tool that raises performance while quietly hollowing out the underlying skill it is sitting on top of.
A separate slice of the cohort, roughly 20%, fell into a category the researchers labeled “Dependency Developers” - participants whose behavior shifted from actively evaluating the evidence to passively accepting whatever the chatbot said.
The MIT News coverage of the study notes that comparable effects have been observed in medicine. An August 2025 paper in The Lancet Gastroenterology & Hepatology found that, after three months of AI-assisted colonoscopy, experienced endoscopists’ adenoma detection rate fell from roughly 28% to 22% when the AI was removed - a pattern that maps directly onto what the Media Lab team saw with news.
Socratic-Style Chatbots Partially Reverse the Effect
There is a partial fix in the data, and it is not “use a better chatbot.” The researchers tested two interaction styles: “telling” chatbots, which hand over a direct verdict (real or fake, with a one-paragraph justification), and “asking” chatbots, which probe the user with Socratic questions before committing to an answer.
Co-lead author Valdemar Danry put it bluntly: “AIs that ‘tell’ by providing direct answers are more likely to foster reliance, while those that ‘ask’ via Socratic questioning are better at engaging someone to actually learn how to discern the truth on their own.” The gain came with a real cost: the Socratic condition was noticeably slower, and the team describes it as “very much a trade-off between speed and effort.”
That trade-off matters because most of the demand for AI fact-checking lives in places where speed matters - a link in a group chat, a headline you scroll past in two seconds, a forwarded image at the dinner table. The interaction style that protects the skill is not the interaction style people actually want.
What This Means
The study is small, four weeks is short, and a 67-person sample is not a population estimate. The authors do not claim otherwise. The result still lines up with a wider pattern that researchers have been naming for a year now: cognitive offloading.
A 2025 Science paper by Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson surveyed 319 knowledge workers using generative AI and found that the more confident a worker was in AI’s ability to perform a task, the less critical-thinking effort they put in. Higher confidence in the AI’s output quality produced more checking; higher confidence in the AI’s competence produced less. The Media Lab study is, in a sense, a controlled experimental version of that finding, with a measurable skill at the other end.
For readers who use chatbots to check news - and that is now a large share of young adults, per Pew - the practical takeaway is uncomfortable. The fastest way to check a claim is also the slowest way to keep the skill. Switching the interaction style from “tell me whether this is real” to “ask me questions so I can work it out” is slower per session but, in the study, partially protected against the dependency effect.
For the people building chatbots, the design choice is sharper. A “tell” mode is the default for product reasons - it is what users reach for, it is fast, it is satisfying. The MIT data is now a published, peer-reviewed argument that this default has a long-term cost the user is not seeing. Designers who want to be honest about that cost have a research-backed reason to ship an “ask me questions” mode as the default for high-stakes checking tasks, even if it loses the speed race.
For schools and newsrooms, the implication is that “let the students use AI to fact-check” or “let the newsroom use AI to triage tips” may be productivity-positive in the short run and skill-negative over a year or two. The same dual result has been seen in medical AI, and Brookings researchers flagged the same “cognitive atrophy” pattern in students earlier this year. The honest framing is to ask what skill the workflow is asking humans to keep, and to design the AI’s role around that skill instead of around throughput.
The Bottom Line
Four weeks of leaning on a chatbot made 67 study participants 15 percentage points worse at spotting fake news on their own, even though they had been 21% more accurate with the bot at the peak. A Socratic-style chatbot slowed the erosion but did not stop it. The lesson is not “stop using chatbots” - the immediate accuracy gain is real - it is that the chatbot is paying for today’s answer with tomorrow’s skill, and most interaction patterns are optimized to hide the bill.