Leadership1 publisher3 min readPublished
Instinctools' CEO would build six foundations before adding another AI agent
Alexey Spas says executives have lost count of the agents already running inside their companies, and cites a 2025 MIT estimate that only 5% of AI solutions produce sustained P&L gains. His fix is a platform layer.
The Board Room · Leadership desk

What happened
- Alexey Spas, founder and CEO of the software firm Instinctools, wrote that executives asked how many AI agents run across their company usually answer some version of "Honestly, I've lost count."
- He cites a 2025 MIT estimate that only 5% of AI solutions generate sustained productivity and P&L gains, and says the agents companies do run collectively deliver nowhere near what was expected.
- He names four faults that repeat across isolated agents: instructions followed with no business goal, decisions made on a sliver of context, context lost at each handoff, and systems that only react to events.
- His post also says only one in four business leaders believes AI agents will act as autonomous workers in the short term, without naming the survey behind that figure.
- The six capabilities he says he would not scale without are cross-stack coordination, shared context, orchestration, processes and KPIs redesigned for a mixed workforce, reusable components and built-in cost controls.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- decision Counting the agents already in production takes an afternoon; the six-capability foundation is a multi-quarter program with no cost stated in the post, and treating the two as one initiative lets a diagnostic anyone can run become a platform purchase.
- constraint Because spend caps and live usage monitoring sit on the foundation list, an estate that grew agent by agent has to retrofit metering across every one of them before central cost control means anything.
- exposure A board underwriting an orchestration program on the 5% figure is relying on a number whose definition of a failed deployment it cannot check from this account.
- precedent If "agentic operating system" becomes the procurement frame, agent budgets shift from the teams that won the early use cases to a central owner who controls routing, caps and handoffs.
One number supports most of the post's case. Spas writes that "by a 2025 MIT estimate, only 5% of solutions generate sustained productivity and P&L gains" [4], and cites it without the study's definition of a solution [17]. Taken at face value, that is 19 deployments that do not pay for every one that does [15]. A buyer cannot tell from the figure whether those 19 failed because they were built in isolation, which is Spas's reading, or because they were picked badly before anyone wrote a prompt.
His causal claim is about the gaps between agents. "When every agent is created in isolation, the only thing that scales is agent sprawl," he wrote [6]. The fault he describes most concretely is the goalless one: "Give an agent instructions but no business goal to measure against, and it'll follow them to the letter, well past the point where a person would have paused to check," he wrote [8].
The trade-off is one of sequencing. Keep adding point agents and each lands a visible win in its own lane, while every agent built before shared context exists is an agent to be re-plumbed later. Stop to build the foundation and the wins pause while a central team defines processes, KPIs, connectors and templates [10]. The post does not say what that build costs or how long it takes [16].
Worth noting who is making the diagnosis. Instinctools is a software engineering company focused on AI-powered digital solutions [1], and Spas calls the assembled result an agentic operating system, a shared orchestration and intelligence layer [13]. Four of his six items are platform work. Two are not: smart routing, spend caps and live usage monitoring can be attached to agents that already exist [11], and they are the items a finance function can check against a bill.
Orchestration as he defines it matches each task to a deterministic engine, an LLM or a human, and logs every handoff in a detailed audit trail [12]. That item has the longest payback. An audit trail is worth little in month one and is the only artifact that answers an internal auditor in month twenty. His worked example is a claim from intake to payout, run so that agents share one business-aware view of the request [14].
The decision an operator faces this quarter is narrower than the six-item list. It is whether to fund metering and a shared context model for the agents already in production before approving the next one. On the evidence in this post, the counting and the metering stand on their own; the rest asks a buyer to take the diagnosis of one vendor CEO on trust.
What to watch
- Whether the 2025 MIT estimate's definition of a solution and its denominator get published, which would tell buyers if 5% is a build problem or a pilot-selection problem.
- Whether enterprise AI budgets move from per-team agent spend to a central orchestration line item with its own owner.
- Whether agent platforms ship spend caps and live usage monitoring as defaults rather than as add-ons bought after the estate is in production.