Skip to content

Build1 publisher3 min readPublished

TestMu AI's Agent Assurance grades agents on the files they changed, not the answers they gave

The launch adds a third verdict, Unable to Verify, and keeps it out of the pass rate. That single choice is more interesting than the rest of the product.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • TestMu AI, founded by Asad Khan, Jay Singh and Mayank Bhola, has launched Agent Assurance, extending its software-testing platform to AI agents that can alter files, call tools, access APIs and open pull requests.
  • TestMu AI says Agent Assurance reads an agent's codebase, generates functional, nonfunctional and adversarial scenarios, runs the agent and checks the resulting files, artifacts and tool calls.
  • The approach moves the test away from an agent's final answer or self-reported activity and toward the systems it actually touched.
  • "The story an agent tells about what it did is the weakest evidence available about its actions," Vipul Verma, TestMu AI's senior vice president of group engineering, said in the launch announcement.
  • A test harness may see the final response without being able to establish whether the correct file changed, the approved API was called or an undeclared tool was used.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

TestMu AI has launched Agent Assurance, extending its software-testing platform to AI agents that can alter files, call tools, access APIs and open pull requests [1]. The company says the product reads an agent's codebase, generates functional, nonfunctional and adversarial scenarios, runs the agent, and then checks the resulting files, artifacts and tool calls [2] - which shifts the unit of assessment away from the agent's final answer and its own account of its activity, toward the systems it touched [3].

That is the right target. An agent's transcript is authored by the thing under test. "The story an agent tells about what it did is the weakest evidence available about its actions," said Vipul Verma, TestMu AI's senior vice president of group engineering, in the launch announcement [4]. A harness that reads only the final response may never establish whether the correct file changed, the approved API was called, or an undeclared tool was used [5].

The most consequential design decision is not the effects-based checking, though. TestMu AI's launch materials assign each validation criterion one of three outcomes - Pass, Fail or Unable to Verify - and describe the unverified portion as an assurance gap that is excluded from the pass rate and can be shrunk by making agents more observable [6]. Marking a criterion Unable to Verify keeps the uncertainty visible instead of quietly scoring it as a pass [7]. It also means a pass rate from this system describes only the criteria the harness could actually observe, so two agents reporting the same number can differ substantially in how much of their behavior was ever seen [8]. Operators reading these reports should ask for the gap alongside the score.

The rest is recognisable QA machinery pointed at a nondeterministic subject: regression tests, smoke tests, CI gates and evidence retention applied to agents whose behavior can change between runs and whose failures reach past a chat window [9]. The documented workflow starts from an agent API endpoint plus requirement materials such as prompts, product requirement documents, knowledge bases, PDFs or DOCX files, which the platform uses to generate scenarios [10]. Agents can be invoked by command, HTTP endpoint, MCP server, or a workflow built on a platform such as n8n [11]. Generated scenarios include prompt injection, instruction overrides and tool misuse by default [12]. Reports separate newly failing, newly fixed and flaky scenarios [13], and the CI commands distinguish an agent failure from an environment failure [14] - the difference between a broken agent and a broken runner is where most teams currently lose their afternoons.

This is contested territory. LangSmith already supports pre-deployment and production evaluations, regression testing, and assessments of agent trajectories and tool calls [15]. TestMu AI's stated differentiator is the evidence layer: checking files, artifacts and tool calls against the agent's declared tool interface, and flagging criteria it cannot confirm [16].

Those claims rest on the company's own product materials. The launch includes no customer case study, comparative test or independent audit of the autonomous-agent system [17], and TestMu AI has published no independent results linking its validation outcomes to fewer security incidents or operational failures; an Unable to Verify result shows the harness lacked evidence, not that an agent is safe in production [18].

Watch for the first assurance-gap numbers from a real deployment. A vendor willing to publish what fraction of its criteria it could not verify against agents running across live corporate systems [19] is making a checkable claim; one that only publishes pass rates is not.

The founding team's history is in this exact business: Khan worked as a lead engineer at GlobalLogic before co-founding 360logica, a testing services firm later acquired by Saksoft, and he and Jay Singh started LambdaTest in 2017 as a cloud service for running web and mobile tests across browsers, operating systems and devices [20].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories