Leadership1 publisher3 min readPublished
A supply chain AI executive relocates the 80% failure problem to the boardroom
Chris Burchett of Blue Yonder argues in Forbes that AI projects fail on strategy rather than technology, a diagnosis that would move the fix from the data science team to the executive suite. What he offers is experience, not measurement.
The Board Room · Leadership desk

What happened
- Chris Burchett, senior vice president for AI transformation at Blue Yonder, writes in Forbes that AI failures he has seen were blamed on technology when they often stemmed from weak strategies.
- He cites RAND for the proposition that 80% of AI projects fail, twice the rate of non-IT projects, and PwC for the finding that only 20% of companies capture 74% of AI's economic gains.
- He reports that his company tested frontier and open-source models on supply chain problems and found domain-specific models more effective and cheaper in some key use cases.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- decision Removing handoffs and consolidating roles is not a technical approval, so if the diagnosis holds, the sign-off and the accountability for failure both move up to whoever owns the operating model.
- constraint Making the degree of change an organisation can absorb an explicit input caps the ambition of any pilot before a model is chosen, which is a harder ceiling to raise than model quality.
- exposure Buyers who act on this are weighing a diagnosis written by someone who sells into the category, and the model comparison supporting it is self-reported, so the verification burden stays with them.
- capability If a domain-trained open-source model really does undercut a frontier LLM on a specialised task, procurement gains a credible alternative to quote in frontier model negotiations.
The claim carrying this argument is not the 80% figure but the causal one behind it, and the column does not close the gap between the two. RAND, as Burchett relays it, counts failed projects [4]; the attribution of those failures to weak strategy rests on his own two decades of practice rather than on any measurement he cites [2][18]. The diagnosis therefore has to be judged on the mechanism it proposes, and the mechanism is about scope.
That mechanism is a claim about where value sits. Praveen Neppalli Naga, Uber's CTO, is quoted saying the biggest wins rarely come from automating one task and instead come from rethinking an entire workflow [6]. Burchett is blunter about the price of doing that: the value is in eliminating or reducing steps, gaps, handoffs and even roles, and his worked example has schedulers redeployed as orchestrators guiding scheduling, load booking and carrier selection agents [9]. No data science function can authorise the removal of a role. If the mechanism is right, the binding constraint is decision rights, and that is the part of the leadership framing the evidence does support.
The concentration statistic is doing less work than it looks. If 20% of companies capture 74% of the economic gains, the remaining 80% share 26% [5], which works out to about 3.7 points of gain per point of company share against 0.325, roughly an eleven-fold gap per company [16]. The RAND comparison, meanwhile, implies a failure rate near 40% for non-IT projects [15], which is not a comfortable baseline either. Neither number tells you whether the leaders lead because their executives sequenced better or because they started with the data, the capital and the clean processes.
A skeptic has an obvious reading available: a vendor's senior vice president of AI transformation [1] telling buyers that failures were strategy problems is also telling them the software was fine. Two pieces of the advice cut against that reading. Burchett tells buyers to prove value in the pilot rather than after scaling, and to be skeptical of vendors who promise the value will appear in production [10]. And his own comparison found that in some supply chain use cases a domain-trained open-source model beat generic frontier LLMs on both effectiveness and cost, which he attributes to warehouse and supplier knowledge being scattered across enterprise systems and frontline staff rather than present in general training data [11][12]. That comparison is self-reported and unpublished, so it earns interest rather than belief.
The interesting tension in the essay is one of sequencing. Uber reported shipping AI agents in two weeks, and Burchett wants a timeline uncomfortable enough to force iteration [7], while also insisting on a proof gate before production, alongside testing for errors and malicious inputs [10][19]. What that combination actually decides this quarter is narrow: which pain point you start on [14], and how much workflow change the organisation will absorb [8]. What it decides next quarter is whether you end up with the reusable stack he recommends [13] or with several pilots that each proved something and none of which compose.
What to watch
- Whether RAND or anyone else publishes a breakdown of AI project failures by cause, rather than a headline failure rate.
- Whether Blue Yonder puts numbers behind its claim that domain-trained open-source models beat generic LLMs on cost and effectiveness.
- Whether Uber reports what the two-week agent shipping produced in durable workflow change, as opposed to launch speed.