How AI Triggers the Reverse Dunning-Kruger Effect

Summary: A new Aalto University study shows that when people rely on AI systems like ChatGPT, they tend to overestimate their own performance—regardless of experience. The familiar Dunning-Kruger pattern, where lower-skilled individuals overrate themselves more than experts, disappears. Instead, users who consider themselves more AI-literate were the most overconfident.

The research highlights a growing problem in human–AI interaction: cognitive offloading. Many users accept AI answers without reflecting, double-checking, or engaging in deeper reasoning. The authors argue that AI literacy by itself is not enough; platforms must also encourage metacognition and critical thinking so users can recognize when AI output might be wrong.

Key Facts

  • Reverse Dunning–Kruger: In this study, users who reported higher AI literacy overestimated their abilities more than novices when working with ChatGPT.
  • Cognitive Offloading: Most participants used only a single prompt per question and accepted AI answers without verification or reflection.
  • Metacognition Gap: Current AI tools do little to help users evaluate their reasoning or learn from mistakes, limiting accurate self-monitoring.

Source: Aalto University

Background

Traditionally, the Dunning–Kruger Effect (DKE) describes how people with lower ability often overestimate their skills, while those with greater competence judge themselves more accurately or even underestimate their ability. But this new research shows that pattern changes when people use Large Language Models (LLMs) such as ChatGPT: almost everyone overestimates their performance, and those who think they understand AI best show the largest confidence gap.

This shows a brain.
Many participants copied the problem into the AI and accepted the solution without checking or second-guessing. Credit: Neuroscience News

The study found that, although using ChatGPT improved raw task performance compared with unaided participants, users consistently overestimated how well they had done. Surprisingly, higher self-reported AI literacy correlated with greater overconfidence rather than better self-assessment.

“We expected AI-literate people to be both better at interacting with AI and better at judging how well they performed with it. Instead, the DKE vanished and more AI knowledge brought more overconfidence,” says Professor Robin Welsch.

The results add to concerns about uncritical reliance on AI: overreliance can weaken people’s ability to find reliable information and may lead to skill erosion in the workforce. Even though ChatGPT helped users arrive at better answers in many cases, the uniform overestimation of performance raises risks for decision-making in real-world contexts.

“AI literacy is important, but technical knowledge alone does not guarantee better interaction or reflection with AI systems,” the researchers note. Current tools often fail to promote metacognition—awareness of one’s own thought processes—so users do not learn from errors or question output sufficiently.

Doctoral researcher Daniela da Silva Fernandes adds: “We need interfaces that push people to reflect. Without that, AI can encourage shallow engagement instead of critical thinking.”

The article was published on October 27 in the journal Computers in Human Behavior.

Why a single prompt is usually not enough

The researchers ran two experiments involving roughly 500 participants who completed logical reasoning items adapted from the Law School Admission Test (LSAT). Participants were split into groups that either used AI assistance or worked without AI. After each problem, participants estimated how well they had performed; they received extra compensation for accurate self-evaluation.

Tasks required substantial cognitive effort, and many participants treated ChatGPT as a quick solution rather than a collaborator. The data showed that most users entered the problem once and accepted the AI’s answer without follow-up prompts or verification. This single-interaction pattern reduced opportunities for feedback, learning, and calibration of confidence.

“We call this cognitive offloading: users shift the mental burden to AI and stop engaging deeply,” Welsch explains. The researchers suggest that requiring multiple interactions or prompting users to explain their reasoning could create better feedback loops and improve metacognitive accuracy.

Practical steps for everyday AI use include encouraging systems to ask users to justify their choices, request clarification on ambiguous responses, or prompt users to verify key facts—measures that force deeper engagement and help expose illusions of knowledge.

Key Questions Answered:

Q: What did researchers find about confidence when using AI?

A: Rather than novices being the most overconfident, the study found that AI-literate users displayed the greatest overconfidence, effectively reversing the traditional Dunning–Kruger pattern.

Q: How does this differ from the classic Dunning–Kruger Effect?

A: With AI assistance, the typical relationship between skill and self-assessment changes: blind trust in AI can erode critical thinking, so even more knowledgeable users misjudge their actual performance.

Q: Why is this important for AI use?

A: Everyone, regardless of AI experience, tended to overestimate their success. This shows that experienced users can still be poor judges of performance when they rely on generative AI.

About this AI research news

Author: Sarah Hudson
Source: Aalto University
Contact: Sarah Hudson – Aalto University
Image: Image credited to Neuroscience News

Original Research: Open access. “AI makes you smarter but none the wiser: The disconnect between performance and metacognition” by Robin Welsch et al., Computers in Human Behavior. DOI: 10.1016/j.chb.2025.108779


Abstract

AI makes you smarter but none the wiser: The disconnect between performance and metacognition

Optimizing human–AI interaction requires that users can critically reflect on their performance, yet little is known about how generative AI affects metacognitive judgments. In two large-scale studies, this paper investigates how using AI relates to users’ monitoring accuracy and performance on logical reasoning tasks.

In Study 1 (N = 246), participants used AI to solve 20 LSAT-style reasoning problems. While AI assistance improved average task scores, participants overestimated their performance by about four points. Higher AI literacy was associated with lower metacognitive accuracy—those with more technical AI knowledge were more confident but less precise in judging their performance. Computational modeling showed that the typical Dunning–Kruger effect disappeared under AI use. Study 2 (N = 452) replicated these results.

The findings highlight how AI can level up task performance while simultaneously weakening users’ ability to evaluate their success. The authors discuss implications for designing interactive AI systems that foster accurate self-monitoring, reduce overreliance, and support cognitive performance through better feedback and metacognitive prompts.