Product1 distinct publisher3 min readPublished
The startup fine-tuned two Nvidia Nemotron models to allow, flag or block each action an agent intends, and reports 98% accuracy on a 9,429-trajectory benchmark whose own authors say accuracy is the wrong test.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The team that rolls this out does not experience it as a benchmark. They experience it as a queue of held actions, a rota of people who have to look at them, and one awkward conversation with the engineer whose Tuesday afternoon vanished into a hold that turned out to be fine.
The gap the product aims at is specific. Permissions answer whether an agent may touch a repository; they do not answer whether this particular write fits the task the agent was handed, and reading logs afterwards names the incident only once it has happened, according to Capsule [3]. So the classifier goes into the path, scoring the intended action just before execution while the customer's policy allows, flags or blocks it [2]. The training material included real agent traces plus adversarial examples written to mark where authorized behavior ends, with humans reviewing the set [11].
Two points of error against a 9,429-trajectory benchmark works out to roughly 189 trajectories [1], the sort of arithmetic that makes a 98% headline feel solid. The benchmark's authors are measuring something else. They report that one rule-based guardrail caught most rogue trajectories while more than three-quarters of its alerts fired on benign code written before anything went wrong, and they argue that accuracy and recall miss the point [6]. The number an operator needs is how often the control halts good work, and the figures in Capsule's release, as relayed by SiliconAngle, are accuracy figures.
The speed claim survives contact better. Ten held steps in one chain add at least 0.71 seconds [2], which is noise beside a test run and only starts to matter if you decide to judge every file read.
The comparison table is where the deck gets ahead of the evidence. The 10.9-point margin [3] is over a model Capsule declined to name, with no score breakdown offered [9], while the frontier systems it says it beat, from OpenAI, Anthropic and Google, are named [10]. What is checkable sits elsewhere. Capsule says billions of tokens across millions of agent interactions already pass through the technology, with financial institutions and technology companies among its customers [13]. Phillip Miller, global chief information security officer at H&R Block, said controls of this kind let security teams widen their use of agentic AI while keeping the security, governance and accountability their clients expect [14]. And on launch day the company disclosed prompt injection flaws in Microsoft Copilot Studio and Salesforce Agentforce, both since patched [16], which is a better credential than an internal scoreboard.
Two axes are worth drawing before anyone signs. One is whether the action can be undone. The other is how often the agent takes it. Rare and irreversible, a wire transfer or a force push to main, is where a synchronous hold earns both its milliseconds and its false positives, and where a human should be paged. Frequent and reversible, reading a file or running the test suite, is where blocking buys the least protection for the most resentment, so log and sample there instead. The two mixed cells are the real argument, and the honest way to settle them is a shadow period where the model flags and never blocks, scored on how many of its flags a human would also have stopped.
Ranked by verification strength, evidence, and original report placement.
Capsule Security Ltd. released a detection system built on two Nvidia Nemotron models it fine-tuned itself, which it calls an "AI circuit breaker" for rogue AI agents, available now.
The models judge an agent's intended action in the moment before it executes, and customers can allow, flag or block it in real time, creating a control layer outside the agent.
Permissions and approval workflows constrain what an agent is allowed to touch but cannot establish whether a particular action fits the task it was handed, and monitoring after the fact only catches the problem once the damage is done.
The StepShield benchmark runs monitors against 9,429 code-agent trajectories drawn from real incidents.
StepShield's authors argue that accuracy and recall miss the point; one rule-based guardrail they tested caught most rogue trajectories, but more than three-quarters of its alerts fired on benign code written before anything went wrong.
Capsule did not identify the third-party model or break out its score.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Solar Pro 4 turns model routing into a procurement decision, not a research one1 distinct publisher
security
Washington names industrial-scale distillation, then hands the detection bill to abuse teams1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Every number originates with the seller
Strip out what Capsule asserts and what remains is a launch date, a funding round, a founder list, two patched vulnerabilities and a description of StepShield. The 98%, the 96.9%, the 71 milliseconds, the halved memory and the token volume are all self-reported, and the one comparison that would anchor them names no opponent. SiliconANGLE reports this carefully; careful relay of a single interested party is still a single interested party.
Shipping, with traction denominated in tokens
The product is generally available, which is more than a demo, and the April vulnerability disclosures show the team works on live agent platforms. But adoption is stated as billions of tokens and millions of interactions — units the customer cannot audit and the vendor chooses — while the buyers are a sector label rather than a name. The one named enterprise voice, H&R Block's global CISO, endorses controls of this kind without saying he has bought these.
A high score on a test its own authors distrust
The gap is not that 98% is implausible — it is that StepShield exists partly to argue accuracy is the wrong number, and the release leads with accuracy anyway. The measure the benchmark's authors care about, how often an alert fires on benign code, appears nowhere in Capsule's disclosure. Add a frontier-beating claim whose opponent is unnamed and a latency floor quoted as a best case, and the marketing reaches further than the evidence. What holds the number down is SiliconANGLE printing the objection rather than burying it.
Launch-day incentives, pointing the same direction
A seed-stage company five months past its public launch needs enterprise credibility, and the fastest route to it is a big number against named frontier labs. Nvidia's interest runs parallel: a security product fine-tuned from Nemotron 3 Ultra and fitting on one L40S is a reference story for both the open model family and the GPU. Nobody in this account is positioned to volunteer a false-positive rate.
Confident about the design, not the performance
We would stand behind the mechanism, the availability, the funding history and the benchmark's construction. We would not repeat the accuracy figures without the vendor's name attached to them, and we cannot say how this behaves on a real engineering team's traffic. A second, independent StepShield run — or one precision number — would move this sharply.