Build1 distinct publisher2 min readPublished
A vendor-commissioned survey of enterprise IT buyers finds monitoring bills, not model quality, ending agentic AI projects. The arithmetic behind the forecasts is worse than the headline.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The unattributable bill has a mechanism, and the mechanism explains why the invoice turns up without a breakdown. One agent task is not one request. It emits a top-level trace, then several model calls, retrieval operations, tool calls, retries and loops, and when the agent delegates to another agent, the trace grows a fresh branch [13]. Every model call carries token, latency, cost and provider data, and every tool call writes its own records for arguments, results, status and downstream activity [14]. The identifiers that would let you split the bill by program, things like tool_name, agent_id and trace_id, are the high-cardinality fields that are hardest to aggregate and most expensive to index [15].
Now the part nobody in the survey write-up computed. Compound the reported spend growth for two years and the average observability bill lands near $5.19 million, about 1.64 times today's [18]. Over the same two years, respondents expect telemetry volume to grow 9.5 times on average [11], with 44% putting their own increase somewhere between 6X and 100X [12]. Those two curves only meet if the cost of carrying a unit of telemetry falls to roughly a sixth of its current level, an improvement of about 5.8X in unit economics [19]. That is not a number you get out of a retention setting or a renewal negotiation.
Two caveats belong on the 9.5X. It is a forecast that respondents made about themselves, and the survey was commissioned by Apica [2], whose chief product and technology officer prescribes intervening earlier in the pipeline as the answer [5][17]. The measured figures are more useful anyway. Among the 54% who saw telemetry triple in a year, AI and ML workloads accounted for 43% of that growth [8], which works out to roughly 0.86 of the original baseline volume added by AI alone in twelve months [20]. Those workloads came close to doubling the estimate on their own, without any of the forecast multiples.
Mann's bank could not say what its AI programs cost, so it cancelled some of them, and he describes AI projects as cannibalising ordinary budgets [6]. That is the operative failure. Only 35% of enterprises claim widespread agentic deployment, and close to two-thirds describe themselves as only somewhat prepared to run those systems [16]. Meanwhile 83% rank AI observability as a top priority for the year ahead [10], which in the bank's case meant the observability survived and the agents did not.
Ranked by verification strength, evidence, and original report placement.
59% of organizations have already terminated or delayed an agentic AI deployment due to monitoring costs.
The finding comes from a survey of more than 300 enterprise IT decision-makers in North America and Western Europe, commissioned by Apica and conducted by Omdia/Informa TechTarget.
54% of enterprises have seen telemetry volume triple in the past year, with 43% of that growth coming from AI/ML workloads, by far the largest driver.
Enterprises report spending an average of $3.17 million on observability, growing 28% year over year.
83% rank AI observability as a top priority for the year ahead.
A single agent task can generate a top-level trace, several model calls, retrieval operations, tool calls, retries and loops, and delegation to another agent adds another branch to the trace.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single vendor-commissioned survey, no methodology published
Every quantitative claim traces to one survey commissioned by the vendor whose product is the article's prescribed remedy, reported in a single outlet with no link to the full report, question wording, respondent profile or margin of error. The technical mechanism (trace fan-out, high cardinality) is well specified and self-evidently checkable, and the internal arithmetic is verifiable, which keeps this above the floor; the causal claims (finance cancels deployments, high-stakes agents most affected) rest on one unnamed-customer anecdote and unquantified assertion.
Self-reported survey adoption only
The only adoption signals are self-reported survey aggregates: 35% claiming widespread agentic deployment, 59% reporting a cost-driven kill or delay, and average observability spend of $3.17M. There are no named deployments, product releases, customer references, usage metrics or measured savings from the prescribed pipeline-first approach, so real-world adoption of the remedy is effectively unevidenced while agentic deployment itself is attested only by respondent self-description.
Overstated relative to evidence and internal arithmetic
Positive gap: the framing ('catastrophic', 'skyscraper', 'panic', 9.5X in two years) outruns both the evidentiary base and the survey's own internal consistency. Compounding the reported 28% spend growth for two years yields only about 1.64X spend against the claimed 9.5X volume, which requires unit telemetry cost to fall roughly 5.8X - an implication the article never addresses. The 59% kill-or-delay rate also exceeds the 35% widespread-deployment rate by 24 points without explanation. The underlying mechanism is real and the cost pressure is plausible, so this is inflation of magnitude and urgency rather than fabrication.
Vendor sponsored the research and sells the remedy
The survey was commissioned by Apica, the sole named expert is Apica's chief product and technology officer, and the article's concluding prescription - an upstream, pipeline-first telemetry control layer - is Apica's product category. Legacy collect-ingest-store-index platforms are cast as the failing incumbent model. Sponsorship is disclosed, which is a mitigating factor, but the research question, the anecdote, the diagnosis and the recommended cure all originate with a single commercially interested party.
Moderate-low
Confidence is limited by a one-source, one-publisher cluster with a commercially interested research sponsor and no methodology disclosure, which caps how firmly the statistics can be relied upon. It is not lower because the sponsorship is transparently disclosed, the figures are internally checkable, the arithmetic tests are deterministic, and the described telemetry mechanism is concrete and consistent with how agent frameworks emit traces.
build
The number a graph benchmark won't print: 740 of 744 operations failed at 40 clients1 distinct publisher
build
A refactoring benchmark stops the best agent at 41.2%, and the tests are the story1 distinct publisher
build
Codex can now ask and keep going, which deletes the only checkpoint you were getting for free1 distinct publisher
leadership
You run three technology estates. Governing them as one is how costs escape the budget.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026