Build1 publisherNot yet confirmed elsewhere3 min readPublished
Dated AI forecasts have no scoreboard, and building against one costs weeks you can count
A developer graded 18 expired AI predictions against their own words. The best mark he shows is a C+, and the one he built on took six weeks out of his own roadmap.
The Engineer · Build desk
What happened
- A developer pulled 18 dated AI predictions from 2024 and 2025, each with a primary source and a deadline already past, and graded them.
- Amodei told the Council on Foreign Relations in March 2025 that AI would write 90% of code in three to six months and essentially all of it in twelve; that deadline passed in March.
- Salesforce now reports 3.8 billion agentic work units served rather than a count of agents, and the billion-agent target is graded an F.
- He had scoped agent workflows into his product Whizi in February 2025 on the assumption Altman was right about the year, then killed the work six weeks later.
- The period's largest exit was on nobody's list: Cursor from $100M to $4B ARR, then a $60B SpaceX acquisition closed on August 14, per the post.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Six weeks is close to half a quarter of a solo product's build capacity, and it was spent on someone else's dated claim; the person who made the claim paid nothing for it.
- constraint The blocker was billing, not intelligence: while a multi-step run is priced per run, agent features stay out of reach at flat subscription prices no matter how good the models get.
- precedent Substituting a new metric for a missed one, rather than correcting it, pushes the work of reconstructing denominators onto buyers and anyone grading later.
- exposure A headline forecast that was wrong about the vast majority of programmers gave cover to a narrower contraction: entry-level workers in AI-exposed jobs are the group actually carrying the loss.
The rule he grades by is the useful part: a prediction is scored against what its own words promised, on its own deadline, and "well, eventually" does not count [6]. Apply it to the thing that actually stopped his build and you get a different lesson from the one the grades imply. A single multi-step agent run, costed against a flat monthly subscription, does not clear [c3b]. Sam Altman's January 2025 wording was that the first AI agents might join the workforce in 2025 and materially change the output of companies [7]. That is a capability date with no price attached, and price is the part a small product has to solve first.
Marc Benioff's September 2024 target was one billion agents on Agentforce by the end of 2025 [9]. Salesforce reported roughly 29,000 cumulative Agentforce deals by early 2026 and crossed $1B in annual recurring revenue [16]. That is a shortfall of about 34,000 to one [22], and the ratio flatters the promise, because a deal is not an agent.
Code generation is where the grade depends entirely on whose count you accept. Google's self-reported 75% of new code as of mid-2026 includes every accepted autocomplete, with humans still reviewing before deploy [18]. Taken at face value, that still sits 15 points under the 90% floor and arrives after the twelve-month deadline for the harder half of the claim [23]. Industry-wide estimates in early 2026 clustered at 25% to 41% [19], so the top of the independent range is under half the promise [24]. The split verdict, C+ for the 90% and F for "essentially all" [11], is about as generous as the evidence allows.
The same generosity shows in the C- for Altman [8]. MIT's NANDA project found in August 2025 that 95% of enterprise generative AI pilots produced no measurable P&L impact [14], and Carnegie Mellon staffed a fake company entirely with agents and measured its best model finishing 30.3% of the office tasks [15]. The passing mark rests on a carve-out for coding agents inside software companies [8], which is the population least like the buyer reading the forecast. Eric Schmidt's April 2025 line that the vast majority of programmers would be replaced within a year gets the F it earned, with programmers still employed where they were [1]. Klarna is the one entry that closed its own loop: the February 2024 claim of work equivalent to 700 full-time agents, then the May 2025 reversal, with CEO Sebastian Siemiatkowski saying the focus on efficiency and cost produced lower quality, and humans rehired [10].
Two things limit how hard this scoreboard can be pushed. It is one developer's, and its inputs are numbers published by the parties being graded, including that autocomplete-inclusive denominator [18] and a deal count standing in for an agent count [16]. Even so, the best mark among the entries he shows is a C+ [12]. The transferable output is not the letter grades. It is the habit of asking for the denominator before you accept the date.
What to watch
- Whether Salesforce ever publishes an agent count again, or whether "agentic work units served" becomes the standard unit other vendors adopt.