Product1 distinct publisher3 min readPublished
A study of more than 109,000 incidents found the destruction stayed invisible until the damage showed up. The same blind spot appears on the invoice: the top 1% of runs carried 46% of one vendor's AI spend.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The dashboard the on-call person opens has a per-seat number in it, because AI sessions get budgeted the way SaaS seats were, at a flat per-seat or per-token rate tracked as an average [11]. Revenium's own 90-day sample puts the median code-implementation task at $2.24 [7]. The stray four-day session cost $3,762 [3], roughly 1,680 times that median [2], about 78 cents across each of its 4,819 calls [1]. A mean built from tasks like the first cannot describe an outlier like the second.
The shape of the spend is the mechanism. Of 14,680 runs tracked over 90 days, the top 1% carried 46% of the total and the bottom 90% carried 12% [8]. One percent of 14,680 is about 147 runs [3]. An alarm set on the total, or on the average, is watching the roughly 13,200 runs that were never going to hurt anyone.
The expensive half is also the half nobody reviews. Across 10,005 interactive agentic sessions the bill came to $109,118, while 4,171 automated pipeline tasks came to $6,723 [9], which works out to $10.91 per interactive session against $1.61 per automated task [4], close to seven to one [5]. The pipeline that opens and reviews pull requests goes through change control. The engineer with a chat window open at four in the afternoon does not.
The destruction cases fail for the same reason one layer down. In the nine incidents in the study ZDNet cites, the agents held valid credentials and did not trip standard monitoring until the damage was already visible [1]. A credential cannot tell whether the caller means to run a migration or drop the table. And the recovery path belongs to someone else: across the 109,000-plus incidents analyzed, median resolution time has been roughly flat since 2023, and the most common fix is waiting for another company's engineers [2]. Worth noting who is holding the notebook here. Revenium sells AI spend management and audited its own practices, which is a vendor with an interest in your alarm. Still, the finding that its own developer's laptop billed $3,762 with nothing budgeted is not the kind of thing a vendor volunteers lightly [12].
The 2x2 for Monday. One axis: can this agent change state you cannot restore inside an hour. Other axis: does it have a hard ceiling that only a human can raise. Reversible and capped, let it run unattended. Capped but irreversible, review per action rather than per session, because the ceiling limits the bill and not the blast radius. Reversible and uncapped is where the 11-day loop lived, about $4,273 a day [4][6], and Revenium's line on that class of failure is that each action was rational in isolation but the cumulative cost was not [6]. The fourth box, irreversible and uncapped, is the one the nine wipes were launched from, at least on the reach side [1].
So cap by run instead of by seat, and make the cap refuse the call rather than send mail. The tradeoff is a developer sitting idle while somebody raises a ceiling mid-task, which will happen and will be annoying; set against a mid-level agentic assistant that BakedWith prices at $100 to $500 a month [10], one four-day accident already costs more than seven months at the top of that band [7].
Ranked by verification strength, evidence, and original report placement.
Revenium engineers reported that on May 13 a developer opened an AI coding session on his laptop that stayed open for four days, ran 4,819 calls and cost $3,762; nobody had budgeted for it and no alert fired.
Of the cumulative-cost cases, the report says: "Each action was rational in isolation, but the cumulative cost was not."
The median cost for agent-based work was $2.24 across 557 code-implementation tasks over 90 days, and the most expensive single task came in at $300.97.
Of 14,680 AI runs tracked over 90 days, the top 1% represented 46% of total spend, the top 5% came to 77% of spend, and the bottom 90% of runs were just 12% of AI spending.
Among 10,005 interactive agentic sessions studied the bill came to $109,118, while 4,171 automated software development lifecycle tasks cost $6,723; the automated pull-request pipeline was under 6% of the bill and the other 94% was engineers using AI through the day.
Revenium's engineers said AI sessions are budgeted similarly to SaaS sessions, with a flat per-seat or per-token cost tracked as an average, and that "SaaS cost controls aim at the wrong part of the curve."
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
product
OpenAI's plan to hand everyone a coding agent leaves the hard part to the model1 distinct publisher
invest
The AI deal frame flipped: buy at 15 times revenue, pay with paper marked at 401 distinct publisher
build
DoiT buys Attribute, and AI cost attribution moves from billing tags to the kernel1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Audited in the wallet, asserted at the headline
Two grades of evidence share one story. Revenium's self-audit is granular to the cent — 14,680 runs, a $2.24 median task, a $300.97 outlier, a $3,762 session with a date on it — and a company documenting its own overspend has little motive to exaggerate it. The nine wiped production systems, which supply the headline and the dek, rest on a study ZDNET never names, dates or links, and BakedWith's price bands arrive with no method at all. The checkable material is the material the headline is not about.
Real usage, one organisation deep
There is genuine deployment underneath the rhetoric: an engineering team that went from seven to 28 AI-using developers, 10,005 interactive sessions logged in 90 days, and a customer whose agent infrastructure allegedly ran from $5,000 to $50,000 a month between prototype and staging. But it is one vendor's telemetry plus anonymous customers, and Revenium's own hedge is the honest reading — the team believes the shape of the distribution holds broadly rather than claiming to have shown it.
The headline outruns its footnote
Nine destroyed production systems set the frame; a spend spreadsheet fills the body. The destruction count and the flat-resolution finding are the two least verifiable items in the piece, while the well-documented parts — tail-heavy spend, averages hiding outliers, interactive use dominating the bill — describe a budgeting failure, not a catastrophe. What joins them is the word 'invisible', used first about monitoring and then about invoices, rather than any shared dataset.
The measurer sells the ruler
Every figure that matters was produced by a company whose business is AI spend management, and the conclusion — that per-seat, average-based SaaS controls 'aim at the wrong part of the curve' — doubles as its pitch. The two most alarming numbers, $47,000 in 11 days and a 10x jump into staging, belong to customers anonymised past the point of checking. ZDNET does label Revenium as an AI spending solutions provider and presents the audit as self-examination, which is more disclosure than these vendor studies usually get.
Firm on what Revenium says, loose on whether it generalises
We can be reasonably sure of the reported contents — the quotes are direct, the internal numbers specific enough to be falsified if anyone tried — and much less sure they describe anyone else's estate. One publisher, one primary document, no replication, and a headline whose study remains anonymous. Confidence would move quickly if a second organisation published its own run distribution.