Leadership1 publisher3 min readPublished
Fortune 50 AI assistant lead traces post-pilot failure to data and ownership
Amirtha Saminathan, who led an AI assistant build at Fortune 50 scale, blames post-pilot failure on data and ownership, citing MIT's 95% zero-return finding. Leaders approving one now choose when to pay for production standards, in a slower pilot or a later rebuild.
The Board Room · Leadership desk

What happened
- Amirtha Saminathan, a data and analytics leader, wrote on forbes.com that the language model behaved fine throughout the customer-facing AI assistant deployment she led.
- She cites MIT's 2025 NANDA study, which found 95% of organizations get zero return on generative AI and put the difference down to approach.
- Her advice to leaders approving an assistant is to build for the production environment from day one, even if the early version looks less impressive.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint An assistant is only as current as its slowest-refreshing source, so money spent on model accuracy cannot buy fresher answers.
- exposure In regulated businesses, answers that cannot be traced or restricted to what a user is cleared to see create compliance risk, and a budget sized to a demo leaves that work unfunded.
- cost The organizational lag is paid in months: finished technology waits while engineering, finance, compliance and operations agree on what success means before usage can be judged.
The version that fits on a board slide is the MIT figure Saminathan cites: 95% of organizations are getting zero return on generative AI, with the study putting the divide down to approach over model quality or regulation [2]. By subtraction, about one organization in twenty is getting some return [3]. The figure counts failures without locating them, and Saminathan's deployment locates one. She had braced for trouble from the language model, and it "behaved fine the whole way through," she wrote [6]. The pilot ran on clean data in a controlled setup with barely any dependencies [7]. Once connected to real systems, the assistant drew on a dozen sources refreshing on different schedules, some by the minute and some once a day, and answered from data that was already hours stale [8].
Her diagnosis is about the organization. "The assistant gets built like a feature when it needs to be built like a system, and the company was never ready to run it once the pilot ended," she wrote [4]. By her account, getting engineering, finance, compliance and operations to agree on what "working" meant was harder than any of the technical work [11]. "The tech was ready months before the organization was, and that gap decided whether the whole thing was worth the spend," she wrote [12]. Her list of decisions to settle before launch opens with ownership: "Who watches performance? Who answers for it when it's wrong?" [16]
I think the ownership questions come first in that sequence. A stale feed is an engineering problem with a known fix, but somebody has to be watching performance to notice the answers are hours old [8][16].
The trade-off she names runs between the pilot and the rebuild. Her advice is to build for the production environment from day one, "even if it looks less impressive early" [14]. The alternative, as she put it: "If you treat the first version as a throwaway and plan to add security, monitoring and real data standards later, then you're not scaling anything, you're just rebuilding it from scratch" [13]. A pilot approved this quarter on the strength of a demo sets up that rebuild in a later one. The rebuild then has to add the identity checks, audit logging and transaction limits the demo never needed [9]. At millions of customer conversations, she wrote, the mismatch "starts costing real money" [5]. She does not say how much [5].
The evidence is one practitioner's account of an unnamed deployment, published on forbes.com [1][6]. Her most portable claim is also the easiest to check. Whether an assistant's sources refresh on compatible schedules is a question an architecture review can answer before launch [8]. The MIT conclusion points the same way as her account [2]. Her broader claims rest on her own experience: that almost none of the value was in the conversation and nearly all of it in integrations with systems of record [10], and that most of what gets called an AI problem is "a data problem or an ownership problem wearing a costume" [15].
What to watch
- Whether MIT's NANDA researchers publish a breakdown of failure causes that separates data and ownership problems from model shortcomings.
- Cost figures from Saminathan's deployment, if disclosed, would show what the post-pilot mismatch costs at millions of conversations.
- Whether other large deployments report the same post-pilot pattern of multi-source feeds going stale once connected to production systems.