Leadership1 publisher2 min readPublished
Aleks data shows ChatGPT-era students 25% less likely to solve word problems without AI
UC Irvine researchers mining a decade of Aleks math data found students 25% less likely to solve word problems in proctored exams after ChatGPT arrived. For teams using AI assistants, it suggests output figures can improve while unaided skill slips.
The Board Room · Leadership desk

What happened
- Working with McGraw Hill, UC Irvine researchers analyzed millions of Aleks sessions from fifth grade through college, 2015 to 2025, using ChatGPT's late-2022 launch as the breakpoint.
- They split the platform's math items into word problems that paste easily into a chatbot and interactive graphing tasks that are hard to paste.
- After ChatGPT launched, students spent sharply less time on word problems and got more of them right, a pattern the authors said showed some students were using AI to solve the problems for them.
- On proctored exams, performance on the graphing problems held steady across the same period.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- exposure Teams judged only on assisted output can lose unaided skill while their throughput numbers improve, leaving the loss off any dashboard until the tool is taken away.
- decision Leaders deploying assistants have to decide whether to test staff without the tool from time to time, since that was the only condition in which the study detected the decline.
- cost If the pattern carries over to work, the loss lands on text tasks that are easy to paste into a chatbot, and it is paid when someone has to do them unaided.
- capability Tutor-style use was linked to better scores at Middlebury, so how an assistant is deployed is a design choice a manager controls, with measured effects on student learning.
For a manager, the useful detail in the UC Irvine study, as Entrepreneur reported it, is where the decline showed up. On the platform, with chatbots available, the measures a supervisor would track got better [3]. The drop appeared only in proctored exams with the tool removed [4], and only on problems that could be pasted into a chatbot [2]. A general slide in math ability should have hit the graphing tasks as well, and it did not [5].
Sina Rismanchian, the U.C. Irvine Ph.D. candidate who led the study, told The New York Times that AI lets students offload the step-by-step thinking that math problems require [6]. Speed comes the week the tool does. Weaker understanding builds over years, and the Aleks record holds about three years of post-ChatGPT data against roughly seven years before it [1].
The nearest this record comes to evidence about work is a separate multi-institutional experiment on adults, cited by Rismanchian and his co-authors. As participants grew used to letting language models produce the answers and skipped the "productive struggle" of working through each step, they became more likely to abandon hard questions they could still solve [7]. In a team, I would expect that to show up as escalations on problems the person could have handled.
Adam Green, a cognitive neuroscientist at Georgetown University, described the pull toward the shortcut in evolutionary terms. "Whatever you could do that exerted the least calories but didn't kill you was the best thing to do evolutionarily," he told the Times [9]. He also told the paper that "Part of how AI is making us dumber" is "undercutting the development of learning how to think in younger people that now have an alternative" [8]. Junior hires learning a role with an assistant already open are the workplace version of the group he describes.
Middlebury College's work suggests the pattern of use changes the outcome. Its July study of more than 200 undergraduates found that students who used AI like a tutor had higher test scores and higher-quality essays than students who had it draft entire essays [10]. Entrepreneur's account does not report how large the gap was.
The decision this quarter is what to measure. A team scored only on assisted output can post better numbers while unaided skill erodes, the same pattern the Aleks homework showed [3][4]. The consequence comes quarters or years later, when the tool is unavailable or wrong. In the Aleks data, the loss was found only by testing without the tool [4].
What to watch
- Fuller methods from the UC Irvine team showing how the pre-ChatGPT comparison cohorts behind the 25% figure were matched.
- A study testing employees who use assistants daily on unassisted work, the setting the Aleks data does not reach.
- Whether McGraw Hill changes how Aleks presents word problems or adds tutor-style AI in response to the findings.