Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Gremlin's Foresight AI agents wait for an engineer before every test and fix

Gremlin has made Foresight AI generally available, an add-on whose agents plan chaos tests, run them, propose fixes and rerun the original test. Every test and every fix still needs an engineer's sign-off, so human review moves to an approval queue.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Gremlin's Foresight AI agents wait for an engineer before every test and fix
Generated illustration
Gremlin says its agents won't test or fix without approval How Gremlin, Wix and Cloudflare each describe the role of AI agents in testing, remediation and deployment.

Gremlin says Foresight AI will neither run a test nor apply a remediation without approval. Wix has described AI interpreting logs and suggesting remediations, with deployment kept under human control. Cloudflare's proposal envisages agents managing more of testing, deployment and maintenance.

Gremlin says its agents won't test or fix without approval
WhoHowKindClaim
Gremlin Foresight AIGremlin says it will neither run a test nor apply a remediation without approvaldecision7
WixHas described AI interpreting CI/CD logs and suggesting remediations, with deployment decisions kept under human controlprecedent10
CloudflareAgent Development Lifecycle proposal envisages agents managing more of testing, deployment and maintenanceprecedent11

What happened

  • Recommendations draw on Gremlin's proprietary Failure Atlas, which the company says spans millions of fault-injection experiments across tens of thousands of systems.
  • Gremlin says system data does not train a language model, and the LLM is limited to searching, summarising and explaining results.
  • A fourth agent, styled a Technical Program Manager, tracks coverage, risks and commitments across teams and sends weekly reports through Slack.

Why it matters

  • constraint Per-item approval caps throughput at the rate engineers can sign off, the same human review capacity Gremlin says AI-assisted delivery is already outrunning.
  • exposure Recommendation quality rests on how closely a buyer's failure modes match a corpus outsiders cannot inspect, so only a trial on real services tests the fit.
  • decision Teams adopting agent-driven testing have to pick where the human boundary sits: before each test as Gremlin does, at deployment as Wix does, or further out as Cloudflare proposes.

The part worth copying is the loop. When a test exposes a weakness, the Operator agent proposes a change and the service reruns the original test [5][6]. The proof that a fix worked is the experiment that failed, now passing. Andrus said: "An alert going quiet doesn't prove anything. A test that passes under load, in realistic conditions, does." [8] We think rerunning the failed test is the right standard. A quiet alert shows a symptom stopped. A passing fault injection shows the failure mode is handled under the conditions that test models.

The fourth agent, a Technical Program Manager, tracks test coverage, reliability risks and commitments across teams, then writes the weekly report and posts Slack updates [5]. Few engineers will miss doing that job. The service also builds dashboards from natural-language requests, meant to show changes in reliability scores and the risks teams have addressed [9].

The second sound decision is where the language model sits. According to Gremlin, the agents decide from platform data and the Failure Atlas, and the LLM handles search, summaries and explanations [16]. That keeps a text generator out of the step that chooses which fault to inject. Gremlin also says system data is not used to train a model [16].

Treat the Failure Atlas like a benchmark table. Gremlin describes it as a proprietary record of millions of fault-injection experiments run over more than ten years across tens of thousands of distributed systems [15]. Its recommendations transfer to a service only if that service fails the way the corpus did, with comparable dependency graphs and comparable retry and timeout behaviour. Because the record is proprietary, a buyer can only test the fit by running it against their own services [15].

Gremlin says Foresight AI will neither run a test nor apply a remediation without approval [7]. It positions the product against AI-assisted delivery, with Andrus arguing that more code is reaching production without close human review [4]. In InfoQ's account he does not attach a figure to that [4]. So a product aimed at a review shortfall still asks a person to approve every test and every fix. The engineer moves from writing experiments to reviewing them. We'd expect a review to cost less than authoring one, but the approval queue still grows with each service enrolled.

Wix has described using AI in CI/CD for log interpretation and remediation suggestions, with deployment and critical infrastructure decisions left to humans [10]. Gremlin's gate covers running the test as well as applying the fix, and we think it belongs there for a first GA release of a tool whose job is injecting failures [7]. Cloudflare's Agent Development Lifecycle proposal has agents managing more of testing, deployment and maintenance [11].

Gremlin's own claim about expertise is narrow. The product is meant to provide reliability checks without every developer becoming a distributed-systems specialist [3]. Engineer Sandip Bhattacharya argued in a LinkedIn discussion that organisations still need specialists to configure agents [12]. Intellyx analyst Jason English, in an article written for Gremlin, cautioned that a reliability programme cannot simply be installed as a software package and needs shared ownership and accountability across engineering and management [13]. Intellyx discloses that Gremlin is its customer [14].

What to watch

  • Whether Gremlin adds a mode that lets agents run tests or apply fixes without per-item approval, moving toward the autonomy in Cloudflare's lifecycle proposal.
  • Customer reports, from sources Gremlin does not pay, on how often Operator-proposed fixes pass the rerun test.
  • Whether Gremlin opens any part of the Failure Atlas to outside inspection.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence30
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Gremlin announced the general availability of Foresight AI, an agentic product that analyses services for potential failures, recommends changes and reruns tests to check that a fix works.

    ReportedSupportedView cited source
  2. [2]

    Foresight AI is available as an add-on to the Gremlin platform.

    ReportedSupportedView cited source
  3. [3]

    Foresight AI is intended to provide reliability checks without requiring every developer to become a distributed systems specialist.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. infoq.com

    1 article · October 11, 2026

    Foresight AI Brings Gremlin Agents to Reliability Engineering

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories