Leadership1 publisher3 min readPublished
Projected 63% rise in AI model spending puts focus on cost per successful outcome
Worldwide end-user spending on AI models and platforms is projected to rise 63% in 2026, and a startup CTO writing for Forbes argues the number a board can audit is what one completed piece of work costs.
The Board Room · Leadership desk

What happened
- Worldwide end-user spending on AI models and platforms is projected to increase 63% in 2026 against 2025, according to a Forbes Tech Council column by Larridin co-founder Ameya Kanitkar.
- The same column cites Gartner's prediction that 60% of organizations will adopt smaller software engineering teams by 2029.
- Kanitkar argues enterprises are building portfolios of models that match capability and cost to the value of each task instead of standardising on one universal AI vendor.
- His proposed purchasing measure is cost per successful outcome, counting model costs, employee time, latency, retries, human review, error rates and the value of the completed work.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- cost The bill for a cheap wrong answer arrives as staff time: extra prompting and hours of human review are counted as part of the price of the task, and they land on the teams doing the checking rather than on the model line.
- constraint Ranking models this way requires logging retries, error rates and review time per workflow. A company can justify running only as many models as its evaluation capacity can check.
- decision The hardware question comes first. Capital committed to one model architecture removes the workload moves the portfolio approach depends on, and the column allows one exception: an environment that must be isolated for security, regulatory or operational reasons.
Kanitkar's two numbers cover different years. The spending projection covers 2026, the year most finance teams are budgeting now [1]. The Gartner prediction about smaller engineering teams is for 2029, three budget years further out [2][15]. A board approving model spend this quarter is answering the first question, and it can revisit the second twice before it falls due.
Kanitkar wrote that "A model that costs half as much per token but requires three attempts is not necessarily less expensive" [6]. Three attempts at half the unit price cost 1.5 times one correct call at full price, so 50% more before anyone prices the checking [13]. He argues the checking is where the money goes: a cheap model that needs extra prompting or creates hours of human review can cost more than a premium model that completes the task correctly the first time [16].
Cost per successful outcome is an operations measurement. Six of the seven inputs Kanitkar counts have to be captured inside the company; the model bill is the only one that arrives from a vendor [14]. A buyer that does not already log retries and review time per workflow cannot compute the figure for one model, and cannot rank three.
The framework is also the product. Kanitkar co-founded Larridin, a Bay Area startup building an organizational platform powered by AI [3]. An argument for spreading work across several models supports a market for the layer that decides which model gets the work. The column credits Gartner for the smaller-teams prediction and does not say where the 63% projection comes from [17]. The narrower claim a buyer can test on its own data is whether a token price predicts what a finished task costs. He stated the rule this way: "The objective should be to use the least expensive model that is capable of reliably delivering the required outcome" [5].
On infrastructure, Kanitkar wrote that "Given the pace of model development, buying dedicated hardware for most AI workloads can create unnecessary rigidity" [9]. He allows one exception: the completely isolated environment a company needs for security, regulatory or operational reasons. Otherwise he puts the work on hyperscalers or specialized AI clouds, which give access to open models without the company buying and managing GPUs [10]. A firm that has already committed capital to one architecture has fewer moves available when a cheaper model appears months later [9].
Open-source models are the other lever he names, both to cut cost and to reduce dependence on proprietary providers [11]. As public tests saturate, he wrote, more teams are evaluating models on their own tasks and data [12]. Many companies still prefer competitive U.S.-developed models over Chinese alternatives for security, governance and procurement reasons [11].
What to watch
- Publication of measured cost-per-outcome figures by an enterprise running two or more models on the same workflow.
- Revisions to the 63% spending projection for 2026 as actual end-user spending is reported.
- Whether hyperscaler and AI-cloud hosting of open models gets priced in enterprise contracts as a substitute for frontier API calls.