Build1 distinct publisher3 min readUpdated
The company's CTO puts agent cost in terminal noise rather than model pricing. The 30 to 80 percent savings are vendor-reported; the 847-to-9 measurement is the part worth keeping.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Yoav Landman, co-founder and chief technology officer of JFrog, told Lets Data Science that in one measured test a coding agent was handed an 847-line log when 9 lines were what it needed to diagnose the failure [1] [2]. That ratio, not the price of a token, is where his account puts the cost of agentic development, and it turns context hygiene into something a finance team can audit [2] [13].
The backdrop is not contested. Per-token prices have fallen roughly 98 percent, with GPT-4-equivalent performance dropping from about $20 per million tokens in late 2022 to around $0.40, a decline of about 50 times [3] [4]. Over the same period enterprise AI bills rose an estimated 320 percent, with average annual budgets going from $1.2 million in 2024 to $7 million in 2026, roughly 5.8 times, and per-developer consumption up about 18.6 times in nine months, according to reporting by The Next Web citing Jellyfish [5] [6]. Uber exhausted its entire 2026 AI coding budget by April [7]. Microsoft revoked its developers' Claude Code licences six months after granting them, with individual engineers spending between $500 and $2,000 a month on tokens [8].
Landman's explanation of the mechanism is mechanical. "Because these agents operate in continuous loops, they retain prior conversation history to preserve their reasoning chain, meaning any terminal noise compounds with every single turn," he told Lets Data Science [9]. Run an npm install or a docker build inside a loop and the window fills with progress bars and status updates that you pay for again on every subsequent step [10]. He puts agentic coding at a thousand times the token consumption of standard code chat, a company figure the publication said it could not independently verify [11].
The product answer has two parts, by his description. JFrog Boost sits locally as a wrapper applying dynamic filters that strip progress bars, ANSI noise and redundant chat logs, with custom TOML filters available for teams whose internal tooling emits its own formats [12] [14]. For codebase search, a component called BoostGraph intercepts search results and prioritises the most relevant files and snippets rather than letting the agent pull in loosely related code [15].
The savings numbers are company-reported: 30 to 40 percent less token waste per CLI command in typical workflows, reaching 80 percent in cases like the 847-line log, with the answer stating plainly that 80 percent is the ceiling and not the average [13]. Applied to Microsoft's disclosed per-engineer range, a 30 to 40 percent reduction would be worth $150 to $800 a month per seat, if spend scales with tokens [16].
The detail that matters more than the percentage is the escape hatch. Landman says agents "can easily bypass filters to retrieve full raw outputs, or automatically disable filters that prove too aggressive" [17]. That is the difference between a slow correct answer and a fast wrong one, because a filter that removes something load-bearing produces confident failure rather than an error. He calls the emerging layer "Agentic Token Efficiency", a phrase from a vendor selling into it, and argues that "more tokens do not equal better output" [18] [19].
Worth watching: whether the per-command reduction shows up on invoices rather than in benchmark decks, and whether bypass events are logged, since the rate at which agents reach past a filter is the only honest measure of how aggressive it was. The 847-to-9 count is also reproducible in-house. Any team can measure the signal share of its own build output before buying anything to fix it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In one measured test, an AI agent was handed an 847-line log when only 9 lines were needed to diagnose the failure, according to JFrog's own measurement.
Per-token prices have fallen roughly 98 percent, with GPT-4-equivalent performance falling from about $20 per million tokens in late 2022 to around $0.40.
Enterprise AI bills rose an estimated 320 percent, with average annual budgets climbing from $1.2 million in 2024 to $7 million in 2026, and per-developer consumption up about 18.6 times in nine months, according to reporting by The Next Web citing Jellyfish.
Yoav Landman is co-founder and chief technology officer of JFrog, and answered Lets Data Science in written answers on why AI coding agents consume so many tokens and what JFrog Boost does about it.
A fall from $20 per million tokens to $0.40 per million tokens is a decline of about 50 times.
An increase in average annual AI budget from $1.2 million to $7 million is a multiple of about 5.8 times.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor measurement, no published benchmark
The mechanism account is specific and internally coherent, and one concrete artefact exists (847 emitted lines versus 9 diagnostic). But every product-performance number is company-reported and unreplicated, the 1000x consumption figure is explicitly unverified by the publisher, the market cost data arrives second-hand through The Next Web citing Jellyfish, and the load-bearing claim — identical task success rates at reduced context — has no published benchmark. Single publisher, single interview.
No adoption data for the product itself
The cluster contains no deployment counts, named customers, usage metrics, pricing or availability details for JFrog Boost or BoostGraph, so product adoption cannot be scored. The adoption-style observations present concern third parties' token spend and licence decisions — they evidence demand for the problem, not uptake of this solution — and inferring product adoption from them would be a guess.
Vendor claims run ahead of published proof
The quantified savings range and the implied outcome — same success rates, faster runtimes, materially lower cost per session — are asserted by an interested vendor with one internal measurement and no benchmark behind them, and the category label 'Agentic Token Efficiency' is vendor-coined. The gap is moderate rather than severe because the publisher discloses the company-reported status, states the 80 percent figure is a ceiling not an average, notes the vendor volunteered that caveat, flags the unverified 1000x number, and separates the durable practical lesson from the product pitch.
Vendor executive promoting his own product and category
Every performance figure originates with the co-founder and CTO of the company selling the product, in a written Q&A that also introduces the category label under which that product would be bought. The commercial interest is direct and undisclosed only in the sense that no third party checks the numbers; the publisher does state the claims are company-reported and that the branding comes from a vendor with a product in the space.
Mechanism credible, magnitudes unconfirmed
Confidence is moderate: the causal mechanism (retained history re-billing tool output each turn) is verifiable from first principles and the one artefact is specific, so directionally the story is sound. But there is a single publisher, a single interested source, second-hand market data, no benchmark on success-rate parity, and no adoption evidence — leaving the magnitude of the savings and the product's real-world reliability unsettled.
build
Three tools, three spellings of the same glob: agent rules do not port1 distinct publisher
build
Your agent needs the API call, not the API key1 distinct publisher
build
Per-developer environments hit their ceiling the day one engineer ran five agents1 distinct publisher
product
amber's 7mn-euro bet: the AI cost centre has moved upstream of the model1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026