Product1 distinct publisher3 min readPublished
IBM's survey of 2,000 technology chiefs puts full readiness for agent deployment at 11%. The gap that shows up in production is authorization, where a limit written per action never sees the sequence it sits in.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
The settings screen asks for one number. Maximum refund the agent may issue without human approval: 50 dollars. In Madhuri Chandoor's worked example, a user asks for ten of those rather than the single 500-dollar refund that would have gone to a person for review, and each of the ten looks permissible when examined on its own [7][8]. Ten times 50 is 500, so the per-action cap tolerates precisely the amount it was written to stop, and it does so without producing a single policy violation to alert on [15].
That is an arithmetic problem before it is a security problem, and it tells you what the log has to contain. A profile can only see a run if every action carries the session it belongs to and the agent identity that took it. Without both fields, each row is its own story, and the story is always fine. Chandoor, who spent two decades in financial services, says the model she wants is fraud monitoring, where the unit of analysis is the account and the run of transactions rather than the charge in front of you [11].
The uncomfortable part is that the agent's reasoning holds up at every step. "A technical conclusion can appear reasonable within a narrow focus," she says, and the context and downstream impact still need weighing before the action proceeds [9].
Keep the evidence separate from the pitch. Both of her cases are illustrations rather than reported incidents [17], and she founded PromptHalo, which sells inspection of why an action is being performed rather than only what is being performed [6]. The measured material belongs to IBM, and what it measures is how executives feel: 11 percent of 2,000 C-level technology executives called themselves fully prepared for the agent deployment they expect within the year [1][2]. Set that beside the 70 percent who said teams were deploying faster than IT could track, and roughly six times as many people report a visibility problem as report readiness [4][16]. Instrumentation comes before profiling in that order, because there is no behavioral baseline for traffic nobody is recording.
Here is a forcing function you can run against your own inventory of agent actions, two axes, no product needed. First axis: can the action be undone cheaply. Second axis: do ten of them add up to something that none of them is. Reversible and non-additive actions are fine under a per-call ceiling, which is where most existing controls already work honestly. Reversible and additive is the refund quadrant, and it wants a session budget and a rate limit instead of a different item ceiling. Irreversible and non-additive wants a scoped credential and one named human. Irreversible and additive is the quadrant to take back out of the agent's hands until the log can join actions to a session and an identity.
The tradeoff is plain. Reviewing aggregates means holding requests that individually qualify, which is latency for the customer and labour for whoever works the hold queue. That is worth paying in the additive quadrants and wasted everywhere else, which is the sorting job, not the tooling job.
Ranked by verification strength, evidence, and original report placement.
A 2026 IBM study surveyed 2,000 C-level technology executives about AI readiness, visibility and control.
In the IBM survey, 11% of the 2,000 C-level technology executives said they felt fully prepared for the AI-agent deployment expected over the following year.
Two-thirds of CIOs and CTOs in the IBM survey said they were accountable for AI systems they did not fully control.
70% of respondents in the IBM survey said teams were deploying technology faster than IT could track.
Madhuri Chandoor is the founder of PromptHalo, which she describes as an AI security and trust infrastructure company that inspects why a certain action is being performed rather than just what is being performed.
In Chandoor's refund example, an agent may issue refunds up to $50 without human review.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
product
Minimus's wind-down makes the hardened base image a continuity line item1 distinct publisher
product
IT's AI shopping list is inverted: 46.5% want automation, 71% of their AI tools are invisible1 distinct publisher
security
96% confident, 63% breached: the confidence number boards should stop accepting1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one survey, one founder
Everything factual traces to a single Next Web article, and inside it to two unverifiable inputs: an IBM study we are given percentages from but no link or methodology, and a founder describing her own company's premise. The one part that survives independent checking is the arithmetic — ten fifties really do make five hundred — which is why the headline holds up better than the reporting under it.
Agents are shipping; the guardrails are a proposal
The only adoption datum in this story cuts against the remedy: agents are going into production fast enough that 70% of technology chiefs say they cannot keep track, while the behavioral profiles and observability gates Chandoor prescribes appear nowhere as something a named organisation has deployed. PromptHalo itself is a description, not a deployment.
The person naming the gap sells the gap
Overstated, but not wildly. A founder in AI trust infrastructure diagnoses a control problem and prescribes her category, and IBM's numbers arrive pre-framed as proof of a widening gap. Two things pull the other way: The Next Web flags the scenarios as hypothetical, and no product claim, benchmark, or efficacy figure is asserted for PromptHalo — the piece argues for a practice rather than selling a result.
Both voices benefit from the worry
Read the sourcing as a stack and the incentive is legible without any speculation: IBM published the survey and supplied its interpretation as a control gap, and every recommendation in the piece comes from the founder of a company whose stated purpose is inspecting why an agent acted. No security team, auditor, or customer appears to say whether any of it works.
Easy to read, impossible to corroborate
We are confident about what was said and by whom — the piece is unusually clear about which claims are hypothetical — and unconfident about whether any of it is true at scale. With a single publisher and no primary survey document, a second account could revise the numbers without contradicting anything we have.