Skip to content

Leadership1 publisherNot yet confirmed elsewhere2 min readPublished

Amazon scraps its developer AI leaderboard within a month after staff gamed token counts

Amazon scrapped an AI token leaderboard for developers within a month of its May launch, after staff ran unnecessary tasks to inflate their counts. Ranking volume spent compute on make-work and left a usage record that cannot show whether the tools improved anyone's job.

The Board Room · Leadership desk

How we use AISend a correction

Photograph accompanying Amazon scraps its developer AI leaderboard within a month after staff gamed token counts
Photo: indiatimes.com

What happened

  • Betterworks' Caitlin Collins said HR teams that track only whether staff use AI tools, and not whether the tools improve their work, collect the wrong employee data.
  • Qualtrics psychologist Benjamin Granger said employee data should be judged by whether it changes how people work in ways that lift performance; raising survey scores is not the aim.
  • Collins said HR seldom names an owner responsible for acting on the findings, and seldom draws up an action plan that loops the results back into the program.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • cost Each unnecessary task run to climb the ranking consumed compute and produced no work, and Amazon bore that cost for as long as the leaderboard ran.
  • constraint Token counts from the leaderboard period cannot separate genuine adoption from rank-chasing, so they are unusable as a baseline for measuring what the tools later achieve.
  • precedent Other employers that rank AI usage now have a documented case of staff running unnecessary work to climb the table, and a reason to expect the same padding in their own counts.

Amazon's scheme, as hcamag.com described it, had two parts, and they measured different things. The weekly target asked each developer a yes-or-no question. With the floor above 80%, fewer than one in five developers could go a week without AI tools before the target was missed [1][15]. The leaderboard asked how much, ranking every developer's token consumption [2]. A threshold is met once a week. A ranking rewards each additional token. The gaming happened on the leaderboard, where workers ran unnecessary AI tasks to raise their counts [3].

The report does not say how many tokens the gaming consumed or what Amazon spent on them. Caitlin Collins, Program Strategy Director at Betterworks, spoke about the general risk of measuring usage alone [4][5]. "If we're not accurately capturing data that shows people are driving value within their job ... then we're losing the thread on this," she said [6].

A skeptic would defend Amazon on the grounds that usage has to exist before anyone can measure what it produces. That argument covers the weekly floor, which pushed breadth of use [1]. It does not cover the ranking. A table of token counts gave developers a reason to inflate volume, and it recorded nothing about whether their work got better [2][3].

Collins starts from the other end. HR should first name the business outcomes a program exists to drive, she said, then define the behaviors that show progress toward them, then identify the day-to-day activities that move those behaviors [7]. Token consumption sits in the last of those layers, and Amazon's leaderboard ranked it directly [2].

Both people quoted in the report sell measurement. Betterworks makes performance management software, and Benjamin Granger is Chief Workplace Psychologist at Qualtrics, an experience management software company [4][10]. Advice to measure outcomes suits their products. Amazon's own month backs the same advice with separate evidence: the volume measure was gamed and dropped within the month it began [3]. Granger said collecting employee data is no longer the hard part [14]. "You've collected the data, because that part's relatively easy," he said [12].

The order of decisions matters most at board level. For the CEO and CFO, Collins said, HR should report an approximate return on investment, the KPIs a program affects, and whether to expand, continue or stop it [8]. A company that ranks usage this quarter will take that choice to its board next quarter with activity counts as its main evidence. "I wouldn't leave surface-level activity data up for anybody to interpret and move on with," Collins said [9].

What to watch

  • Whether Amazon keeps the 80% weekly-use target on its own; the report says the initiative was scrapped but does not separate the target from the ranking.
  • Any figure from Amazon on the compute consumed by unnecessary tasks during the leaderboard month.
  • Whether Amazon or other large employers begin reporting AI programs to boards on ROI and KPIs in place of token counts.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories