Skip to content

Leadership1 publisher3 min readPublished

Five harness rebuilds in two years anchor Jabali AI's case against waiting for better models

Jabali AI founder Vatsal Bhardwaj argues that the engineering around a model explains the 95% of enterprise AI pilots that MIT found showed no P&L impact. His five rebuilds in two years point to a standing budget, on evidence from one games company.

The Board Room · Leadership desk

Illustration accompanying Five harness rebuilds in two years anchor Jabali AI's case against waiting for better models

What happened

  • Jabali AI founder Vatsal Bhardwaj wrote that the harness around a model, meaning its context, tools, sandboxes, workflows, memory and evaluation, decides whether a demo becomes a product.
  • Bhardwaj cites MIT research finding that 95% of enterprise generative AI pilots deliver no measurable P&L impact.
  • He also cites a Gartner forecast that more than 40% of agentic AI projects will be canceled by the end of 2027.
  • Jabali rebuilt its harness five times in 24 months, moving from a rigid orchestrator through a human-approval copilot to autonomous agents writing code in isolated sandboxes.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision Budget held back for the next model release leaves the feedback and workflow gaps Bhardwaj attributes to MIT untouched, because the deploying company has to build both itself.
  • capability A company that owns its harness can take each better model as a cheaper swap; one without a harness pays for every model upgrade as a new integration.
  • cost On Jabali's evidence, green dashboards can pass a failing product, so a pilot budget has to include paid human time spent reading raw output.

The board-deck version of the stalled pilot fits on one slide: most pilots show no return, the models are not ready yet, so hold the budget until the next release. Bhardwaj wrote that he has "come to believe the comfortable explanation is also the expensive one" [4]. In his words, the distance between a demo and a dependable product is one that "bigger models don't close" [15]. His support is a second finding he attributes to the MIT research: pilots stall because systems don't learn from feedback and don't fit real workflows [6].

A skeptic would say this is one founder generalising from video games. That is an accurate description of the evidence. Jabali builds systems that turn a plain-language idea into a playable game [10], and every lesson in the column comes from that workload. Bhardwaj's answer is that games are a harsher test than text, because "an unplayable game fools nobody" [11]. That makes his account good evidence about how AI failures stay hidden during a pilot. The 95% figure counts pilots without P&L impact; blaming those failures on the harness is his inference from one company [2].

His most useful point is about models, and it complicates his own harness-versus-model framing. The best model for Jabali's workload changed repeatedly, he wrote, and each switch got cheaper because the harness absorbed the change [8]. Better models kept arriving, and Jabali adopted them [8]. "Never marry a model," he wrote [14].

The harness is also recurring work. Five designs in 24 months comes to one rebuild roughly every 4.8 months [12]. This quarter's decision is whether harness engineering gets a standing budget line or a one-off project code. Next quarter's consequence, if Jabali's pace is typical, is that the first design is already due for replacement. The column does not put a cost on any of the five rebuilds.

Jabali's evaluation runs in layers: objective checks such as whether the output runs, then coherence checks, then subjective quality scored by AI judges calibrated against human ratings [7]. Early on, a gruff blacksmith written to grunt a few words about ore would, ten minutes into a playtest, brighten into a cheerful assistant offering numbered quest options [9]. "Nothing errored, and every metric we had was green," Bhardwaj wrote [9]. The team found the fault, which it named persona drift, because someone spent a week reading raw transcripts [9].

Whether the pattern holds for enterprise pilots outside games, we do not know yet. The Gartner cancellation forecast Bhardwaj cites runs to the end of 2027 [3]. That makes this a planning question for next year. The decision this quarter is smaller: which one or two pilots get harness and evaluation staff, so the claim is tested on the company's own workload.

What to watch

  • Whether MIT's research, read directly, ranks feedback and workflow fit above model capability as the cause of stalled pilots.
  • Any enterprise outside games reporting pilot results before and after investing in orchestration and evaluation while holding the model constant.
  • Gartner's next update on agentic AI cancellations ahead of its end-2027 horizon.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories