Skip to content

Leadership1 publisher3 min readPublished

McKinsey meters its consultants' AI tokens like mobile data on a work phone

Consumption pricing made model access a variable cost, and McKinsey now watches it user by user, emailing the heavy spenders rather than cutting them off. EY reports a larger saving from a quieter method.

The Board Room · Leadership desk

Photograph accompanying McKinsey meters its consultants' AI tokens like mobile data on a work phone
Photo: businessinsider.com

What happened

  • McKinsey tracks AI consumption user by user and emails employees when their usage runs high, a system introduced firmwide in the summer, according to QuantumBlack UK lead Debasish Patnaik.
  • By May 2026 the firm was processing about five trillion tokens a month, with roughly 10% of users accounting for about 65% of that, consultants and software engineers among the heaviest.
  • Beyond the alerts, McKinsey runs an internal gateway that optimizes requests before they reach providers and circuit breakers that pause access while it judges whether heavy usage is productive.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint Circuit breakers put a ceiling under the autonomy pitch: access can be paused pending review, which limits how far "we don't cap" describes the arrangement.
  • exposure Because consumption is concentrated in a named tenth of staff, cost control becomes a conversation about identified individuals rather than about workflows or tools.
  • precedent EY's quantified 60% cut sets the disclosure bar for firms selling AI cost governance, and a qualitative account of changed habits now reads as the weaker submission.

The concentration figure alone argues against hard caps. If a tenth of users generate about 65% of tokens [6], each of those users consumes roughly 17 times what an average user in the other nine-tenths does: 65 points of usage split ten ways is 6.5 each, while the remaining 35 points split ninety ways is under 0.4 [19]. A ceiling set anywhere near the average would land almost entirely on the consultants and software engineers the firm names as its heaviest users [6], in the cases Patnaik says still pay for themselves [9].

The autonomy framing is real but narrower than it sounds. Alongside the alerts, McKinsey runs circuit breakers that temporarily pause access around particularly high token usage while the firm assesses whether that usage is productive [11]. That is a limit with a review attached. What the firm has settled, then, is where the limit sits and who has to justify crossing it.

Nudge emails are also what a firm sends when the bill has not started to hurt yet, and the record supports that. Patnaik said internal AI spending has not reached a problem stage, and that usage could be more "egregious" in another six months [9][10]. The reported effect so far is qualitative: employees changed how they used the tools once they could see their own consumption [8]. EY gave a number instead, telling Business Insider that an "invisible" router directing staff to the model best suited to a task, with other governance measures, cut token consumption by 60% since April [13][14]. The two mechanisms differ in who does the work of saving the money. A router saves the money automatically, without needing the user to act; an email relies on the user acting instead.

There is a measurement mismatch inside the advice. Patnaik tells clients to assess token cost per outcome rather than per employee [18], while the instrument McKinsey has built reports per employee, because a gateway sitting between users and model providers is positioned to see users [1][11]. Nothing in the account describes tokens being attributed to engagements or deliverables. A per-user number turns a cost review into a conversation about named people; a per-outcome number would need plumbing no one here claims to have.

The demand for any of this is coming from the buy side. McKinsey set up a formal practice this spring to advise clients on deploying AI cost-effectively [16], and Patnaik said clients spent the last quarter finding a "hidden cost that we haven't completely thought through" [17]. The cause is upstream: providers moved from subscriptions to consumption-based pricing this year [3], and individual ceilings are now high enough that OpenAI said in September its most prolific coding-agent users consume more than $7,000 of tokens a day [4]. When GitHub changed its pricing in June, a senior software engineer at Deloitte US told Business Insider the new monthly quotas were "already wreaking havoc" on what could be expected of developers [15]. Firms shopping for a playbook are shopping for a governance design, and McKinsey's puts the burden of restraint on the people getting the most out of the tool.

What to watch

  • Whether McKinsey ever publishes a spend or token-per-outcome figure to sit alongside EY's reported 60% reduction.
  • Whether the circuit breakers fire often enough to function as quotas, especially if Patnaik's six-month scenario arrives.
  • Whether model and tool providers reprice again, as GitHub's June quota change did for Deloitte's developers.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories