Skip to content

Product1 publisher2 min readPublished

OpenAI's deployment lead blames rollout and trust for 80% of stalled enterprise AI projects

Colin Jarvis, who runs OpenAI's forward deployed engineering team, says deployment is the block for most enterprise AI buyers. The way his engineers are staffed and measured shows what that costs.

The Product Desk · Product desk

Photograph accompanying OpenAI's deployment lead blames rollout and trust for 80% of stalled enterprise AI projects
Photo: thenextweb.com

What happened

  • Colin Jarvis, who leads OpenAI's forward deployed engineers, said about 80% of companies struggling with enterprise AI have a deployment problem: they do not know how to roll it out, govern it or prove they can trust it.
  • In the remaining 20% of cases customers say the model cannot yet do the job, and Jarvis said those gaps are now narrow and specialist, such as a task in semiconductor design.
  • At one semiconductor company, where engineers spent maybe 80% of their time on bugs from overnight jobs, OpenAI built a system that patches them after review, and the customer estimates savings of roughly $40m to $50m a year.
  • Custom work is now about 50% of each project, down from about 90% before tools such as Codex, and hiring has moved toward domain experts including a former investment banker and a chip verification engineer.
  • AWS is spending $1bn on the same approach of putting engineers inside client companies, and Microsoft launched a $2.5bn deployment business in July.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost With about half of each project still bespoke, the growing line item is senior people's time inside the buyer's own building, and a cheaper model per token does not shrink it.
  • constraint If the remaining capability gaps are specialist ones, a team that has parked its office-work automation until the next model release has parked it on a condition that has already been met.
  • exposure A business case that leans on the 80% figure is leaning on one executive's on-stage estimate from the company that also sells the deployment help.
  • precedent Three vendors now fund engineers who sit inside customer companies, so deployment labour becomes a priced part of an AI purchase and procurement has to review it the way it reviews consulting.

Every engagement starts with a two-day visit. The first thing the team asks business leaders to do is ignore AI and name the biggest levers in their business [8]. Whatever comes back is the project [8]. Jarvis said one of the two mistakes he sees most often runs the other way: companies pick a use case because it seems to fit AI [11].

The second mistake is the use case that works in one department and stays a demo [11]. Against that, he described a semiconductor customer with about 35 use cases live after roughly 18 months [13]. It built a central team to scale projects and placed small groups of engineers in each business unit [13]. That pace is about two new production use cases a month, held up for a year and a half [21].

FDEs have no financial incentive tied to usage, Jarvis said, and are measured on whether a project reaches production and moves a real metric [14]. When OpenAI's embeddings were too slow for a Klarna search service, he told the company to use an open-source model instead [15]. "From OpenAI's side, we should always be temporary," he said [16].

The diagnosis comes from OpenAI's head of forward deployed engineering [2], speaking to Alex Hern of The Economist at HumanX in Amsterdam [3]. Jarvis gave no source for the 80% figure. On the safety question, Jarvis was working from memory: he said he thought OpenAI paused its main reinforcement learning run "in September this year" [17], while OpenAI's own post on the pause, published 18 August, describes a two-week pause in reinforcement learning training on its latest models and says the largest planned frontier run stayed on hold [18]. "I don't personally feel a lot of pressure to race ahead on model development itself," Jarvis said [5].

Two tests come out of his account. Both can be run before any model is picked. One: whether a business leader named this lever when nobody had mentioned AI [8]. Two: whether the work has an owner with authority outside the department where it started [13]. If both are true, the work can spread. A named lever with no owner beyond the originating department is the demo Jarvis described [11]. An owner with no named lever gets something into production that moves nothing anyone reports [14]. Fail both tests and you have a proof of concept; the companies that succeed, Jarvis said, measure success by production use [12].

What to watch

  • Whether OpenAI publishes any data behind Jarvis's 80% estimate.
  • Whether the semiconductor customer ever states its own $40m to $50m saving estimate publicly.
  • Whether AWS's and Microsoft's deployment arms show up as separately priced line items in customer contracts.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories