Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Dated AI forecasts have no scoreboard, and building against one costs weeks you can count

A developer graded 18 expired AI predictions against their own words. The best mark he shows is a C+, and the one he built on took six weeks out of his own roadmap.

The Engineer · Build desk

How we use AISend a correction

What happened

  • A developer pulled 18 dated AI predictions from 2024 and 2025, each with a primary source and a deadline already past, and graded them.
  • Amodei told the Council on Foreign Relations in March 2025 that AI would write 90% of code in three to six months and essentially all of it in twelve; that deadline passed in March.
  • Salesforce now reports 3.8 billion agentic work units served rather than a count of agents, and the billion-agent target is graded an F.
  • He had scoped agent workflows into his product Whizi in February 2025 on the assumption Altman was right about the year, then killed the work six weeks later.
  • The period's largest exit was on nobody's list: Cursor from $100M to $4B ARR, then a $60B SpaceX acquisition closed on August 14, per the post.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Six weeks is close to half a quarter of a solo product's build capacity, and it was spent on someone else's dated claim; the person who made the claim paid nothing for it.
  • constraint The blocker was billing, not intelligence: while a multi-step run is priced per run, agent features stay out of reach at flat subscription prices no matter how good the models get.
  • precedent Substituting a new metric for a missed one, rather than correcting it, pushes the work of reconstructing denominators onto buyers and anyone grading later.
  • exposure A headline forecast that was wrong about the vast majority of programmers gave cover to a narrower contraction: entry-level workers in AI-exposed jobs are the group actually carrying the loss.

The rule he grades by is the useful part: a prediction is scored against what its own words promised, on its own deadline, and "well, eventually" does not count [6]. Apply it to the thing that actually stopped his build and you get a different lesson from the one the grades imply. A single multi-step agent run, costed against a flat monthly subscription, does not clear [c3b]. Sam Altman's January 2025 wording was that the first AI agents might join the workforce in 2025 and materially change the output of companies [7]. That is a capability date with no price attached, and price is the part a small product has to solve first.

Marc Benioff's September 2024 target was one billion agents on Agentforce by the end of 2025 [9]. Salesforce reported roughly 29,000 cumulative Agentforce deals by early 2026 and crossed $1B in annual recurring revenue [16]. That is a shortfall of about 34,000 to one [22], and the ratio flatters the promise, because a deal is not an agent.

Code generation is where the grade depends entirely on whose count you accept. Google's self-reported 75% of new code as of mid-2026 includes every accepted autocomplete, with humans still reviewing before deploy [18]. Taken at face value, that still sits 15 points under the 90% floor and arrives after the twelve-month deadline for the harder half of the claim [23]. Industry-wide estimates in early 2026 clustered at 25% to 41% [19], so the top of the independent range is under half the promise [24]. The split verdict, C+ for the 90% and F for "essentially all" [11], is about as generous as the evidence allows.

The same generosity shows in the C- for Altman [8]. MIT's NANDA project found in August 2025 that 95% of enterprise generative AI pilots produced no measurable P&L impact [14], and Carnegie Mellon staffed a fake company entirely with agents and measured its best model finishing 30.3% of the office tasks [15]. The passing mark rests on a carve-out for coding agents inside software companies [8], which is the population least like the buyer reading the forecast. Eric Schmidt's April 2025 line that the vast majority of programmers would be replaced within a year gets the F it earned, with programmers still employed where they were [1]. Klarna is the one entry that closed its own loop: the February 2024 claim of work equivalent to 700 full-time agents, then the May 2025 reversal, with CEO Sebastian Siemiatkowski saying the focus on efficiency and cost produced lower quality, and humans rehired [10].

Two things limit how hard this scoreboard can be pushed. It is one developer's, and its inputs are numbers published by the parties being graded, including that autocomplete-inclusive denominator [18] and a deal count standing in for an agent count [16]. Even so, the best mark among the entries he shows is a C+ [12]. The transferable output is not the letter grades. It is the habit of asking for the denominator before you accept the date.

What to watch

  • Whether Salesforce ever publishes an agent count again, or whether "agentic work units served" becomes the standard unit other vendors adopt.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories