Skip to content

Leadership1 publisher2 min readPublished

Wharton's agency decay scale tracks how a team slides from testing AI to depending on it

Wharton's Scale of Agency Decay runs experimenting, integrating, relying, depending. The account details the first two, and its usable marker is whether people still overrule a system that contradicts them.

The Board Room · Leadership desk

Illustration accompanying Wharton's agency decay scale tracks how a team slides from testing AI to depending on it

What happened

  • Wharton defines agency decay as the gradual erosion of a person's ability and willingness to observe carefully, think independently, choose deliberately and act responsibly.
  • The Scale of Agency Decay has four stages, and Wharton says people and organizations have travelled it faster than their governance, training and culture could follow.
  • A study in The Quarterly Journal of Economics found a generative AI assistant raised customer-support productivity by 15% on average, with the largest gains among less experienced and lower-skilled workers.
  • The Canadian Moffatt v. Air Canada decision held the airline responsible after its chatbot gave a customer misleading information about bereavement fares.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • exposure So far the cost of unverified output has landed on the professional who signed the filing and the company that deployed the bot. Whoever supplied the model has not carried it.
  • constraint A checking step added at stage two is funded out of the same productivity gain that justified buying the tool, so the business case and the control compete for the same minutes.
  • decision Because the stage markers are behavioural, telling integrating from relying means sampling decisions for whether anyone overrode the system, since the outputs will grade well either way.
  • cost The staff with the largest measured gains are the ones with the least experience to evaluate what comes back, and that bill arrives later, in a cohort that cannot check work.

Verification is what Wharton's stage two turns on. The mild form is cognitive offloading, handing memory and routine work to the system [5]. The harder form the account names is belief offloading, where people outsource evaluation as well as effort, and where the operating question moves from "Is this task suitable for AI?" to "What did the AI say?" [6]. Research on automation bias and algorithm appreciation, the article says, finds individuals accepting AI-generated answers as correct even in the presence of contradictory evidence [7].

If a support agent resolves 15% more cases an hour [4], the time per case falls by about 13%, so a ten-minute case becomes roughly 8.7 minutes [16]. That is 1.3 minutes [16]. It is also the whole budget from which a verification step gets funded, unless the team hands back part of the gain that bought the tool.

In Mata v. Avianca, the lawyers who submitted filings containing fabricated cases generated by ChatGPT were the ones sanctioned [10].

The account details experimenting and integrating, and it names relying and depending but never says what separates them [15]. As supplied, the scale is a vocabulary, and any threshold a company sets will be its own. Two behaviours in the text are observable now. Users frequently defer to AI suggestions even when those suggestions conflict with their own judgment, and that deference is stronger when the system has performed well in prior tasks [8]. Wharton also reports that lower trust in humans correlates with higher trust in AI [9].

In enterprise settings, the article says, AI is embedded in customer service, coding environments, legal research and financial analysis, often as the first point of contact [14]. Those are the workflows where a wrong answer reaches a customer, or a court, or a set of accounts. Workers there increasingly treat the systems as collaborators and drop their outputs straight into drafts, analyses and recommendations [18].

The embedding is a this-quarter fact. Wharton dates it three and a half years from ChatGPT's arrival as a research preview on November 30, 2022, by which point AI sat inside search engines, office software, customer-service platforms and coding environments [12][13]. Erosion of judgment is measured over careers, and Wharton says individuals and organizations moved along the scale faster than their governance, training and culture could follow [3]. The choice available this quarter is narrower: which of those workflows keeps a recorded check, and whose name is on it.

What to watch

  • Whether the relying and depending stages get defined with thresholds a manager can observe, or stay as labels.
  • Whether another court extends the Moffatt reasoning from a customer-facing chatbot to an internal AI-assisted decision.
  • Whether replication work finds the 15% support-agent productivity gain holds for experienced staff as well as junior ones.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories