AI raises the score, and overconfidence

People who used ChatGPT to solve hard reasoning problems scored higher than people who solved them without it. But they also believed they had done better than they had.

What the researchers did

Two studies led by Aalto University in Finland, published in Computers in Human Behavior, had participants use ChatGPT to solve 20 logical reasoning problems from the Law School Admission Test, the entry exam for US law schools. After each problem they judged how well they had done, and were paid extra for judging accurately.

In the first study, 246 people used ChatGPT. Their average score was 13.0 out of 20, against 9.5 for a comparison group of 3,543 people who had solved the same problems without AI. On average, they thought they had scored 16.5. A second study, with 452 people, replicated the findings.

 

Average scores out of 20: 9.5 without AI, 13.0 with ChatGPT, and 16.5 for what the ChatGPT users thought they had scored.


On tasks like these, the people who perform worst usually overestimate themselves the most. Whats known as the Dunning-Kruger effect. With AI, that effect ceased to exist. Across the board, people overestimated their performance.

Higher AI literacy made it worse. Participants with more technical knowledge of AI were more confident, and less precise in judging their own performance.

"What's really surprising is that higher AI literacy brings more overconfidence," said Professor Robin Welsch.

Most participants rarely prompted ChatGPT more than once per question. They copied the question in and accepted the answer without checking it. The researchers call this cognitive offloading: the processing is done by the AI.

That shallow engagement may have left people with too few cues to judge how well they had actually done. Read alongside the researchers' discussion, the findings suggest higher AI literacy makes higher cognitive offloading and lower engagement in the solution finding process more likely, underscoring concerns about workforce de-skilling.

"Current AI tools are not enough. They are not fostering metacognition and we are not learning about our mistakes," said doctoral researcher Daniela da Silva Fernandes. Her suggestion is for AI tools to ask users to explain their reasoning, so that they engage with the problem and face what they do not yet understand.


What this means at work

Judging our own capability accurately is hard at the best of times. Colleagues and managers tend to agree with each other about a person more than either agrees with the person themselves. AI-assisted work adds a second layer: a better result, and a less accurate sense of how good it is.

Four questions to ask where AI is part of the work:

  1. When a piece of work was produced with AI, how would the person know if it was wrong?
  2. Can they explain the reasoning without the tool open?
  3. Who, besides the person who produced it, checks AI-assisted work before it is relied on?
  4. Are people building the skill, or handing it over each time the task comes up?


Source

Fernandes, D., Villa, S., Nicholls, S., Haavisto, O., Buschek, D., Schmidt, A., Kosch, T., Shen, C. & Welsch, R. 2026. Performance and Metacognition Disconnect when Reasoning in Human-AI Interaction. Computers in Human Behavior, 175, 108779. doi.org/10.1016/j.chb.2025.108779

Aalto University. 2025. AI use makes us overestimate our cognitive performance. aalto.fi/en/news/ai-use-makes-us-overestimate-our-cognitive-performance

Back to blog

Be the first to know

New products, tools, and resources from Edaith — straight to your inbox.