Skip to content

Leadership1 publisher3 min readPublished

Acumatica's CEO pins AI's unprovable ROI on a missing pre-pilot baseline

Writing for Forbes, John Case argues that AI's proof problem is a measurement problem, citing MIT research in which 95% of studied organizations saw no P&L impact six months after a pilot. His company sells the systems that would hold the baseline.

The Board Room · Leadership desk

Photograph accompanying Acumatica's CEO pins AI's unprovable ROI on a missing pre-pilot baseline
Photo: diginomica.com

What happened

  • John Case, chief executive of cloud ERP company Acumatica, wrote in Forbes that the harder question he hears from businesses starting with AI is how to quantify its impact on the business.
  • His worked example is a finance team that closes the books faster after adopting AI, where nobody can say how many hours were saved because the process was never measured beforehand.
  • Case wrote that some Acumatica customers told him their employees went back to old processes because the AI was too kludgy or poorly integrated into their workflows.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint Without a measured before-state, the evidence that would justify or kill the spend cannot be reconstructed later, so renewal decisions get made on how the process feels to the people running it.
  • contradiction A board that reads the 95% and the two-thirds as one trend is stacking a financial test on a perception survey, and will price task-level gains as if they had already reached the income statement.
  • decision If the baseline has to live in existing business systems, the measurement question turns into a systems-consolidation question, and the vendor making that argument sells those systems.
  • exposure Reversion puts the risk on projects already booked as wins: value recorded at pilot can leak back out once staff quietly return to the old workflow.

The two figures Case cites come from different instruments. MIT's is a financial test: 95% of the organizations studied showed no measurable profit-and-loss impact from generative AI investments six months after a pilot [3]. Upwork's is a survey of perception: more than two-thirds of small and midsized business leaders reported productivity gains, and most of those gains were below 25% [4]. Subtract the MIT figure from the whole and roughly one in twenty studied organizations did register something on the P&L [1]. The column does not name the MIT paper or give either study's sample size or date [16].

A 20% gain on a task and a flat income statement can both be true in the same quarter. Case's own list of useful AI outcomes opens with cost reduction, meaning lower outside spend, fewer manual processing costs or fewer corrections [15]. Freed hours become one of those only when someone takes the cost out. He allows that a 10% or 20% gain is still meaningful, while noting that expectations often fall short of the dramatic transformation associated with AI [5].

The baseline argument is the part that survives the vendor's framing. Case describes a finance team that brings in AI to close the books faster; six months on everyone agrees the process feels smoother, but nobody can say how many employee hours were saved or whether error rates moved, because those were never measured first [6]. "Businesses can also mistake activity for outcomes," he wrote [7]. Adoption rates and hours of usage indicate engagement and say much less about whether the business improved [8].

Case writes that most SMBs cannot build a new measurement infrastructure for every experiment, and that "Their existing business systems have to carry much of that load" [10]. He is CEO of Acumatica, a cloud ERP company, after nearly 30 years in cloud services [1]. He also identifies disconnected customer, financial, inventory and operational data as a reason companies spend hours reconciling before they can evaluate anything [9]. The diagnosis points at the product he sells.

A skeptic would say six months is early for any technology program to reach a P&L, and that the 95% figure is a stopwatch problem. Case's own treatment of time answers that from the other side: early gains may show up as shorter cycle times, greater employee capacity or better access to information, and leaders still need to see whether those improvements repeat [12]. His durability evidence is anecdotal: some Acumatica customers told him their employees reverted to old processes because the AI was too kludgy or poorly integrated into their workflows [11].

For anyone approving spend this quarter, the consequence is narrow and unglamorous. A pilot funded now without a measured before-state can only be defended in two quarters by the testimony of the people who ran it. Case's prescription is to answer two questions before deploying, what are we trying to improve and how are we measuring it today, and to start with a manageable workflow such as customer service, administrative work, analytics or inventory, where the link between the technology and the result is relatively clear [13][14].

What to watch

  • Whether MIT's and Upwork's underlying studies are published with sample sizes, dates and populations a CFO can check against the 95% and 25% figures.
  • Whether measurement at 12 or 18 months, rather than six, moves the share of pilots with a traceable P&L effect.
  • Whether ERP and business-system vendors start selling pre-pilot baselining as a priced module.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories