Skip to content

Leadership1 publisher2 min readPublished

Fragmented AI architecture, not just costly inference, is why enterprises struggle to prove AI ROI

Gartner found 72% of CIOs at break-even or worse on AI and expects more than 40% of agentic projects to be canceled by 2027 on cost. A platform founder argues both figures show that every new agent rebuilds the same plumbing.

The Board Room · Leadership desk

Illustration accompanying Fragmented AI architecture, not just costly inference, is why enterprises struggle to prove AI ROI

What happened

  • Gartner reported that 72% of CIOs say their organizations are breaking even or losing money on their AI investments.
  • Gartner also forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027 because of escalating costs.
  • Ragy Thomas, chairman and co-CEO of the integration platform UnifyApps, argues the cost comes from bolting AI onto fragmented architecture, with every new agent rebuilding context, integrations, controls and verification.
  • He argues the same fragmentation makes ROI harder to prove, because costs, decisions and outcomes stay scattered across separate systems.
  • On a shared platform, he writes, each workflow leaves one record of the data used, the actions taken, the human intervention required, the cost incurred and the outcome delivered.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • cost McKinsey's finding on retrofitted governance means the bill is paid twice, and the rework and delay land on whoever shipped first without controls.
  • decision Defining controls once turns a platform question into a scheduling question: the second agent either inherits reusable controls or sets the house rule that each team writes its own.
  • exposure Ten separately built claims agents mean ten places where access to policyholder data can be configured wrong, and an auditor or regulator reaches all ten.

When every step of a workflow runs on the same model, the simple steps pay for reasoning they never use [7]. Researchers from the University of Hong Kong and Stellaris AI studied one financial-services enterprise whose inference costs passed $200,000 a month, while more than 70% of its queries were routine enough for smaller models [8]. That is a run rate above $2.4m a year [9]. The case counts queries, not dollars, so how much of that bill the routine traffic consumed is a separate question [8].

The swarm argument runs the same logic one layer up. A set of narrow agents avoids reprocessing the full context at every step, because each one holds a purpose-built context window and passes forward only the output the next agent needs, which Thomas says cuts token cost per task and response time [10]. The efficiency holds only if the swarm is governed as one system, with shared permissions, controls and observability across agents [11].

Governance is the piece that gets rebuilt: a production claims workflow needs role-based access to policyholder data, audit trails, limits on autonomous actions and escalation rules [17]. Build ten agents independently and much of that framework gets designed, tested, documented and monitored ten times [12]. Define the controls once and apply them as code across every workflow and integration, and compliance overhead grows more slowly because the controls are reused [13]. Gartner's estimate is that effective governance technologies "could reduce regulatory expenses by 20%" [15].

Thomas sells the platform he is diagnosing. He is chairman and co-CEO of UnifyApps and founder and chairman of Sprinklr [1], and he wrote: "CIOs are continually asking me the same two questions: What is AI actually costing us, and is it delivering measurable ROI?" [6] The two halves of his argument have different provenance. Both figures above are Gartner's, and Gartner's stated cause for the cancellations is escalating cost [2][3]; the step from cost to fragmented architecture is his [4]. The 72% also groups break-even with loss, so on Gartner's own count 28% of CIOs report doing better than break-even [16].

The tradeoff is who pays for reuse. A shared layer for context, integrations and controls costs the first workflow schedule time, and the tenth workflow inherits it. Thomas's own example uses ten agents [12]; his piece does not say at what agent count the reuse covers the platform. Routing work to smaller models and splitting it across narrow agents both work by buying fewer tokens [7][10].

What to watch

  • Whether Gartner revises the 40% cancellation forecast as 2027 approaches, and whether escalating cost remains the stated reason.
  • Published before-and-after token cost per task from a governed swarm deployment, reported by someone other than a platform vendor.
  • Whether the University of Hong Kong and Stellaris AI work breaks the $200,000 monthly bill down by query class, which would show how much routing could actually recover.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories