Skip to content

Build1 publisher3 min readPublished

IBM's skill-erosion survey records worry from 60% of 8,800 employees

IBM and Oxford Economics found that 60% of 8,800 full-time employees worried AI was eroding their skills. The closest controlled test, a 52-developer Anthropic experiment, supports building practice into AI-assisted work and is too small to blame the tools for skill loss.

The Engineer · Build desk

Illustration accompanying IBM's skill-erosion survey records worry from 60% of 8,800 employees
Generated illustration

What happened

  • IBM announced the results on September 21 from surveys run with Oxford Economics from April through June that also covered 1,500 human-resources leaders.
  • A Microsoft Research survey of 319 knowledge workers, presented at CHI 2025, linked higher confidence in AI with less self-reported critical thinking.
  • The dev.to post relaying the findings proposes giving each AI-assisted task two finish lines, one for the working feature and one for the developer's understanding.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A count of worried employees cannot size actual skill loss, so a training plan justified by the 60% figure still needs its own test of what people can do.
  • cost On the Anthropic evidence, unstructured assistant use while learning a new library costs comprehension and buys no significant time saving, which weighs against it for onboarding.
  • decision Team leads who want both assistant speed and retained skill have to set a checkable learning target per task, since a passing feature shows nothing about whether its author understood it.

A count of worried employees and a test of what employees can do measure different things. The IBM figure is the first kind. Its numbers come via a dev.to post summarising IBM's announcement [1]. Sixty percent of 8,800 respondents is roughly 5,280 people reporting a concern [3][2]. The post's author wrote that it is "a report of concern, not a controlled measurement proving that AI caused people to lose ability" [4].

The post points to an Anthropic experiment published in January as closer evidence. Anthropic randomly assigned 52 developers learning an unfamiliar Python library to work with or without an AI assistant [5]. The assisted group averaged 50% on a quiz afterward. The unassisted group averaged 67% [6]. The gap is 17 points, about a quarter of the control group's score [6][1]. Completion time did not differ significantly between the groups [7].

The result belongs to a specific workload. The post describes the study as small, covering one learning task, with the quiz shortly afterward, and says it does not establish long-term effects on developers generally [8]. Its comparison of interaction styles was observational, so it cannot show that one prompting technique caused better learning [9]. For the number to transfer, your people have to be learning a library they do not know and be checked soon after. A senior engineer working in a codebase she knows well is outside the conditions tested.

Inside those conditions, the timing result matters more than the score. The assisted developers scored lower and did not finish significantly faster [6][7]. For onboarding, where the learning is the deliverable, that study offers no time saving to set against the lost comprehension.

The Microsoft Research study is the weakest of the three on causation. It surveyed 319 knowledge workers for a CHI 2025 paper and found higher confidence in AI associated with less reported critical thinking [10]. Those were self-reports and associations [11]. The direction could run either way: people who already check less may simply trust the tool more.

The evidence supports treating this as a training design problem. The post proposes a design at the level of a single task: "give the feature and your understanding separate finish lines" [12]. Its worked example is a task tracker that sorts by due date. The developer writes down the intended order first. Then they predict where tasks with an early date, a later date and no date will land, and what breaks a tie between two matching dates, before the assistant shows the result [13]. The suggested prompt tells the assistant: "Ask me to predict the order for a small example before showing the answer" [14].

In my view this is the right unit for a team lead running onboarding. It costs minutes per task. It also leaves a written prediction that a reviewer can compare with what the code actually did. The author is careful about its standing, calling it "a work habit you can try, not a scientifically validated recipe" [15].

What to watch

  • IBM's full survey release, where question wording and breakdowns by role would show what respondents meant by skills eroding.
  • A larger replication of the Anthropic experiment with a delayed retest, to show whether the 17-point gap persists beyond a quiz taken shortly after the task.
  • A controlled comparison of predict-before-reveal prompting against plain assistant use, the habit the dev.to post proposes without validation.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories