Skip to content

Product1 publisher3 min readPublished

Employers are scoring workers on AI usage in performance reviews

Leaderboards at Disney, Meta, JP Morgan and KPMG rank how much staff use AI, and Accenture's chief says promotion depends on working that way, which leaves reviewers holding a usage count rather than evidence the work got better.

The Product Desk · Product desk

What happened

  • Accenture chief executive Julie Sweet told the Rapid Response podcast in March that AI is how the firm does work now, and that promotion requires doing the things Accenture does to operate.
  • Disney, Meta, JP Morgan and KPMG have set up AI leaderboards that track and rank employees' usage of the LLMs and platforms available to them, according to media reports cited by the BBC.
  • Coinbase has already fired engineers who did not complete the AI training its chief executive Brian Armstrong asked for.
  • Duncan Trevithick, a 34-year-old marketer at an AI training data company, qualifies for a year-end bonus if he can use AI to finish more tasks faster.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint The only column a reviewer can export is visibility, so the person who ships good work without touching the tools sits in the same bucket as the person coasting, and the ranking cannot tell them apart.
  • cost The time AI frees up accrues to the employer as a reset baseline, and what the employee gets back is a chance at promotion rather than the hours or the equivalent pay.
  • exposure Long tenure stops working as cover: the senior person with no visible AI use is the one who can now be overtaken by a three-year hire who is quick with the tools, on Pamela's account.
  • decision With the legality settled, according to Chander, the choice left to each manager is which artifact the review cites when someone is passed over, and whether it survives being read aloud.

A leaderboard's rank order is the cheap part. It counts how often someone opened which model [6]. Whether the deck or the pull request that came out was worth shipping sits somewhere else, usually in a manager's recollection of it.

A review needs two axes, not one. One axis is visible tool use, and the leaderboard supplies it. The other is work the team would defend to a client, and the leaderboard does not. High use plus good work is the person the scheme was written for. High use plus thin work is the person it rewards by accident, because the only column carrying a number is the one they lead. Low use plus good work is the person it marks down: the executive the BBC calls Pamela, a US-based senior figure at a large consultancy, says people who are not visibly using AI see slower promotion timelines and find it harder to be seen [11][13]. Low use plus thin work is the one box nobody argues about. None of the reporting shows how any of these employers scores the second axis.

McKinsey puts 94% of companies as yet to see significant value from AI [8], which leaves 6% that have [1]. A firm sitting in the 94% is measuring inputs because the outputs have not arrived, and a usage export is available this quarter in a way that a value case is not.

Duncan Trevithick's arithmetic shows the trade plainly. He says AI handles about two days of work a week for him, and that he receives neither the two days off nor a 40% pay rise [3]. Two days of a five-day week is 40% of the week, so the raise he names is his freed time priced at par [2]. What he gets instead is a shot at promotion and a higher baseline to be measured against next cycle [3].

Tina Chander, an employment lawyer and partner at Weightmans, told the BBC that employers doing this are legally in the clear [14]. Her open question is the fairness one: whether it is reasonable to raise expectations because the employee can now produce more, and whether the more efficient employee should be rewarded differently or is simply making themselves redundant [15]. Lawful and defensible are separate tests, and the second is the one a manager fails out loud in a calibration meeting.

The hiring market looks calm about it. Of the 1,881 UK jobseekers Gi Group polled in July, 75% said AI proficiency in individual performance reviews would not put them off applying and around 22% said it would [10], leaving roughly 3% who said neither [3]. The BBC sets that against UK vacancies at a five-year low [9]. That tolerance says more about a thin job market than about acceptance of the metric.

The sentence a reviewer would have to say to the person passed over is the real test, before AI proficiency ever goes into the template. If it comes out as "your usage was in the bottom quartile", the criterion is measuring the tool, and the employee is entitled to ask who the gain went to. If it comes out as "here is what you shipped and here is what it took you", the tool never needed to be in the criterion. Trevithick describes his own job now as managing some people and also managing AI agents, or loops, or whatever you want to call it [4]. That is the version a reviewer can defend, because it names what somebody took responsibility for and what it produced.

What to watch

  • Whether any of the named leaderboard employers publishes a written definition of AI proficiency tied to output rather than usage counts.
  • Whether McKinsey's 94%-see-no-significant-value figure moves in the next survey round, which would change whether firms are measuring inputs for want of outputs.
  • The first grievance or tribunal claim in which a usage ranking is the evidence behind a bonus, promotion or dismissal decision.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories