Build1 publisher3 min readPublished
Sonnet's output rate caps a $20 seat at 1.3 million tokens a month
The multipliers in circulation for how far AI seats are underpriced run from five times to twenty, and all of them rest on an unpublished per-seat token count. The API rate card is the part you can check.
The Engineer · Build desk

What happened
- A dev.to analysis argues the $20 Claude Pro and ChatGPT Plus seat is being served at roughly five times the cost the lab collects for it, and that the arrangement is not stable.
- Claude Code sessions run autonomously, reading files, writing files and running commands, and users have exhausted five-hour rate-limit windows in under ninety minutes.
- GitHub moves Copilot to usage-based billing on June 1, 2026, with the announcement citing agentic usage becoming the default and a qualitatively different inference demand.
- The KPMG Q1 2026 pulse has U.S. organizations projecting average AI spending of $207 million over the next twelve months, roughly double the year before.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction Five times, eight times and twenty times are three different budget outcomes, and one piece supports all three, so a buyer cannot size a future reprice from the figures now in circulation.
- decision Seats are budgeted by headcount and metered bills are charged by token, so any team that wants a forecast has to start counting output tokens per seat before the pricing changes underneath it.
- precedent With one vendor metering agentic use, the rest have cover to follow, and renewal conversations move from how many seats a company needs to how many tokens those seats burn.
- exposure Marketing copy and code review were wired through subsidised seats over two years, so a reprice lands on teams that did not pick the vendor and cannot see the meter.
The one figure anyone can check here is the rate card. Sonnet is $15 per million output tokens on the API [3]. At that rate, $20 buys about 1.33 million output tokens a month, and that is a ceiling, since it assumes the seat pays nothing for input [17]. Spread across 22 working days, it comes to roughly 60,000 output tokens a day [18]. Ordinary chat fits inside that: the dev.to piece puts a normal session at a few thousand tokens and heavy use in the tens of thousands [13]. A developer running three or four coding agents in parallel moves close to ten times the tokens of the same person in chat [15].
The multipliers in circulation do not agree with each other. The piece puts the gap at about five times [1], then prices an equivalent knowledge-worker seat at $200 to $400 a month [4], which against a $20 seat is ten to twenty times [19]. The Anthropic analysis it cites, roughly $8 of compute for every $1 of subscription revenue, lands at eight [6][21]. Each of those numbers is a function of tokens per seat per month, and the piece does not state the token volume behind its $200 to $400 estimate [22]. For that range to transfer to your seats, your users would have to move the tokens its author assumed.
The labs already have a throttle in place. Exhausting a five-hour allowance in ninety minutes means consuming at about 3.3 times the rate the window was sized for [20]. A rate limit cuts what a seat effectively gets without any change to the sticker price.
OpenAI's language has moved. Nick Turley, the company's VP of product, described the subscription pricing as something they "stumbled into", and has floated phasing out unlimited plans, comparing them to "unlimited electricity" [8]. Sam Altman said the company now needs to become "an AI inference company" [9].
The buyer side has started counting. A Goldman Sachs survey of large companies found most of them overrunning their AI budgets by orders of magnitude [11]. Chandrasekaran, who runs AI and data at KPMG North America, told Marketplace: "Even a quarter or two ago nobody bothered about LLM consumption costs" [12]. Microsoft was reportedly losing more than $20 a month on every GitHub Copilot seat, with power users reaching $80 [5].
In my view the number a team can defend at a budget review is its own: output tokens per seat per day, measured. With it, a metered reprice is a multiplication you can run in advance. Without it, the seat line tells you how many people have logins. The $20 sticker has not moved since 2022, while the models behind it picked up image generation, code execution, voice, agentic reasoning and web search [7].
What to watch
- Whether Anthropic or OpenAI publishes per-seat token allowances, which would let buyers compute the multiplier for their own users instead of inheriting someone else's estimate.
- The first metered Copilot invoices after the billing change, and whether heavy users land anywhere near the $80 figure reported for the flat-fee era.
- KPMG's next pulse survey, and whether the $207 million projection tracks what organizations actually spend.