Build1 distinct publisher3 min readUpdated
A developer graded 18 expired AI predictions against their own words. The best mark he shows is a C+, and the one he built on took six weeks out of his own roadmap.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The rule he grades by is the useful part: a prediction is scored against what its own words promised, on its own deadline, and "well, eventually" does not count [4]. Apply it to the thing that actually stopped his build and you get a different lesson from the one the grades imply. A single multi-step agent run, costed against a flat monthly subscription, does not clear [c3b]. Sam Altman's January 2025 wording was that the first AI agents might join the workforce in 2025 and materially change the output of companies [5]. That is a capability date with no price attached, and price is the part a small product has to solve first.
Marc Benioff's September 2024 target was one billion agents on Agentforce by the end of 2025 [9]. Salesforce reported roughly 29,000 cumulative Agentforce deals by early 2026 and crossed $1B in annual recurring revenue [10]. That is a shortfall of about 34,000 to one [1], and the ratio flatters the promise, because a deal is not an agent.
Code generation is where the grade depends entirely on whose count you accept. Google's self-reported 75% of new code as of mid-2026 includes every accepted autocomplete, with humans still reviewing before deploy [14]. Taken at face value, that still sits 15 points under the 90% floor and arrives after the twelve-month deadline for the harder half of the claim [3]. Industry-wide estimates in early 2026 clustered at 25% to 41% [15], so the top of the independent range is under half the promise [4]. The split verdict, C+ for the 90% and F for "essentially all" [16], is about as generous as the evidence allows.
The same generosity shows in the C- for Altman [8]. MIT's NANDA project found in August 2025 that 95% of enterprise generative AI pilots produced no measurable P&L impact [6], and Carnegie Mellon staffed a fake company entirely with agents and measured its best model finishing 30.3% of the office tasks [7]. The passing mark rests on a carve-out for coding agents inside software companies [8], which is the population least like the buyer reading the forecast. Eric Schmidt's April 2025 line that the vast majority of programmers would be replaced within a year gets the F it earned, with programmers still employed where they were [17]. Klarna is the one entry that closed its own loop: the February 2024 claim of work equivalent to 700 full-time agents, then the May 2025 reversal, with CEO Sebastian Siemiatkowski saying the focus on efficiency and cost produced lower quality, and humans rehired [13].
Two things limit how hard this scoreboard can be pushed. It is one developer's, and its inputs are numbers published by the parties being graded, including that autocomplete-inclusive denominator [14] and a deal count standing in for an agent count [10]. Even so, the best mark among the entries he shows is a C+ [5]. The transferable output is not the letter grades. It is the habit of asking for the denominator before you accept the date.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Eric Schmidt said in April 2025 that "in the next one year, the vast majority of programmers will be replaced by AI programmers"; the year is up, programmers remain employed almost everywhere they were before, and the prediction is graded F.
The author pulled every dated 2024 and 2025 AI prediction he could find with a primary source and a deadline that has already passed; eighteen made the cut and were graded.
In March 2025 Dario Amodei told the Council on Foreign Relations that AI would be writing 90% of code within three to six months and essentially all of it within twelve months; that twelve-month deadline passed in March.
In February 2025 the author scoped agent workflows into his product Whizi on the assumption Altman was right about 2025, and killed the project six weeks later.
He killed the agent project after doing the math on what a single multi-step agent run costs against a flat monthly subscription.
The grading rule: a prediction is graded against what its own words promised, on its own deadline; if you said 90% by September and September brought 30%, "well, eventually" is not a grade.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-published post, mostly uncited relays
The cluster is a single dev.to article by an individual developer. Its strongest evidence is internal and checkable in principle - dated verbatim quotes from named executives, a stated grading rule, and a first-person account of a killed project. Everything that carries the argument's weight externally (NANDA's 95%, CMU's 30.3%, Stanford/ADP's 16%, Salesforce deal and work-unit counts, Google's 75%, and a $60B Cursor acquisition) is relayed second-hand with no links, dates, or filings, and no second publisher in the cluster corroborates any of it.
Real deployment, far below the forecast line
The supplied material does contain adoption signal, and it points the same way in every instance: Agentforce reportedly at roughly 29,000 deals and $1B ARR against a billion-agent promise, AI code share at a caveated 75% inside Google and 25-41% industry-wide against a promised 90%, enterprise pilots mostly showing no P&L impact, an agent-staffed company finishing 30.3% of tasks, and Klarna re-hiring humans after an automation claim. Adoption is genuine and commercially non-trivial but sits far below the graded predictions, and each figure comes from one uncorroborated relay.
Debunk directionally aligned, its own numbers oversold
The forecasts being graded are clearly overstated relative to the adoption evidence presented, and on that axis the article is aligned with reality rather than adding hype. The residual positive gap belongs to the article itself: a punchy scoreboard framing ('the scoreboard is brutal', a ~34,000-to-one shortfall computed by comparing promised agents to reported deals, a record-breaking acquisition claim) rests on single-source figures with no citations, so the piece asserts more precision than the supplied evidence can carry.
Practitioner promotion plus contrarian-format pull
The author writes on a developer publishing platform and uses his own product, Whizi, as the narrative spine, which creates a visibility incentive; the accountability-scoreboard format also rewards maximally quotable grades and shortfall ratios. On the other side, the graded parties - vendor CEOs announcing agent, coding, and workforce timelines - had direct promotional and fundraising incentives to publish aggressive dated claims, which the article documents through their own words. No compensation, sponsorship, or vendor relationship is disclosed in the supplied material.
Low - single truncated source, uncorroborated figures
Confidence is limited by structure rather than by internal contradiction: one publisher, one author, a body truncated mid-sentence, and the central external statistics unlinked. The consistently repeated pattern across independent-sounding examples raises confidence in the direction of the finding, while several specific figures - the acquisition, the deal count, the payroll study - would each need a primary source before being treated as fact.
invest
The 81% Problem: AI's Star CEOs Are Polling Badly With The People They Need To Hire1 distinct publisher
product
Nine AI leaders, nine majorities of distrust: the floor onboarding copy cannot lift1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
product
Wu says Cognition is not for sale. The more useful fact is who bought Cursor last week.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 22, 2026