Leadership1 publisher3 min readPublished
An AI pilot stalls where two teams define "active customer" differently
Shawn Rosemarin of Pure Storage argues in Forbes that different numbers for one metric mean the definitions were never aligned, and that the expensive part of fixing it is refinement work most organizations have never scoped.
The Board Room · Leadership desk

What happened
- Shawn Rosemarin argues in a Forbes Tech Council post that when two leaders bring different numbers for the same metric both are often correct, and that "the definitions were never aligned."
- His example is a single word: one team counts a customer as active on a recent transaction, another counts an open account as enough.
- He also cites Gartner's prediction that 60% of AI projects will be abandoned through 2026.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- decision The choice sitting with leadership is which reported metrics get one owner and one written definition, because no further platform spend settles a disagreement about meaning.
- cost Refinement scales with the volume you choose to refine, so savings from consolidating storage do not cover the work that comes after it. That work lands on data and domain teams, and infrastructure budgets do not carry it.
- constraint If meaning is established at capture, a repair applied at the reporting layer keeps reopening. That reopening caps what another dashboard rebuild can deliver.
- exposure A business case built on the 60% abandonment figure is exposed on causation: the forecast covers abandonment, and the attribution to unrefined data comes from the essay's author.
Definitions are set upstream, at capture, where data is labeled. The middle stage in Rosemarin's scheme is the one most consolidation projects fund: move it, store it, optimize for volume and throughput instead of usability. That middle stage is what he calls the "death zone" for AI projects [9][10]. Two leaders arrive with different numbers drawn from the same three-year platform build, both technically correct. The argument is about what meaning was attached at the source and which definition each team applied at the end [2].
The spend follows from that. Rosemarin writes that raw storage is affordable and that refinement is the expensive part, scaling with how much you decide to refine [14]. That makes the scope of refinement a budget line, and until it has an owner it is set by whoever loads the most data.
The cost varies by data type. In his grading, structured tables and CRM records are light crude; semi-structured logs and nested files are heavy crude. Photos, video, call recordings and email are oil sands, abundant and far more expensive to recover [11]. Before a model can use a contract, he writes, something has to read it, link it to a customer identity and attach permissions [12]. Shammy Narayanan, senior vice president for data, AI and architecture at Welldoc, told Rosemarin that over half of clinical data in healthcare is stranded behind technical debt. The bytes, he said, are already paid for [13].
The two figures in the essay come from Gartner, cited second-hand. A 60% abandonment rate through 2026 works out to three of every five funded projects expected to be dropped [7][17]. The 63% figure combines data leaders who lack foundational practices with those who are unsure whether they have them, and the essay does not break the two groups apart [6][18]. Gartner supplies the forecast; the cause comes from Rosemarin, who writes that the reason is unrefined data and not model failure [8].
Rosemarin is a vendor executive writing a council post. His account is built from meetings he has sat in. A diagnosis that clears the platform is a comfortable message from someone whose day job is customer engineering at a storage company [1][19]. The claim is also cheap to check without buying anything: one metric, two teams, and a comparison of the definitions each used.
The two repairs run on different clocks. Naming a single owner and a single definition for each metric leadership reports on is available this quarter and costs meeting time. Reading contracts, linking identities and attaching permissions across unstructured data is years of work. Rosemarin's harder claim is that some of it should never start: where recovery and refinement cost exceeds expected value, leave it in the ground [15]. He writes that most organizations have never made that decision explicitly [16].
What to watch
- Whether Gartner publishes a breakdown separating data leaders who lack foundational practices from those merely unsure they have them.
- Whether any large buyer discloses a refinement budget line distinct from storage and platform spend.
- Whether healthcare systems put a number on the clinical data Narayanan says is stranded behind technical debt.