Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Gremlin's Foresight AI reruns the failing chaos test to check its own proposed fixes

Gremlin released Foresight AI on October 7th to propose fixes for the faults its chaos tests expose. The launch came with no beta accuracy or acceptance figures, so reliability teams have only their own supervised trials to judge whether those fixes hold.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Gremlin's Foresight AI reruns the failing chaos test to check its own proposed fixes
Generated illustration

What happened

  • Gremlin's homepage says Foresight AI proposes tests and changes for the team to approve, keeping a human sign-off step in the loop.
  • Diagnoses draw on the test, service and health-check context, with each recommendation tied to observed test results, according to Gremlin's documentation.
  • A proprietary Failure Atlas of more than a decade of cause-and-effect data on failures and recovery underpins the recommendations, Gremlin says.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Reliability teams can adopt Foresight AI under the approval gate and use their own accept, edit and reject counts as the first accuracy evidence they can check.
  • exposure Granting the external-model permission lets service context reach an outside AI service, so the data-flow review has to clear before anyone judges diagnosis quality.
  • cost Buyers carry the work of confirming how to halt a running test and revert an applied change before any proposed patch reaches a production system.

Foresight AI starts from a chaos experiment that has already exposed a weakness. Gremlin positions it against AI tools that respond to incidents already underway, aiming instead at failure modes a team can test in advance [10]. Going by Gremlin's materials, a run has four steps:

1. A controlled fault test exposes a weakness in a service [2]. 2. The product diagnoses it from the test, service and health-check context, tying its recommendations to the observed results [6]. 3. It proposes a configuration patch or an infrastructure-as-code change for the team to approve [2][8]. 4. It reruns the test that exposed the risk, and can repeat tests as the system evolves [2][11].

Step four is the good engineering here. Runtimewire put the case plainly: a recommendation can be plausible and still leave the failure mode untouched, and rerunning the originating test checks the result under the same condition [13]. A passing rerun covers one injected fault. Side effects of the change on paths that test never exercised have to surface in the rest of the suite.

Step two draws on what Gremlin calls its Failure Atlas, a proprietary set of more than a decade of cause-and-effect data on failures and recovery [5]. Treat it like any benchmark built on someone else's workload. For its diagnoses to transfer, the failures in that data have to resemble the dependencies and recovery behaviour of the stack under test.

The model question is separate. Use of an external large language model is optional and requires customer permission, Gremlin says [7]. Runtimewire notes that buyers are able to check which system context the product draws on and under what conditions it may be passed to an AI service [14]. Granting that permission is a different decision from approving any single change.

Gremlin says Foresight AI completed a beta. The announcement also leaves out the beta's customer count, its duration and any incident-prevention results [3][9]. Runtimewire concluded that Gremlin's claim to have already proven the product can identify and address risks before incidents is unverified by those measures [4].

I think the right deployment is a supervised pilot on non-critical services. Keep the approval gate Gremlin describes on every change [8], and log each outcome as accepted, edited or rejected. Counted over the pilot, those outcomes give an acceptance rate measured on the team's own services.

What to watch

  • Gremlin publishing beta figures: customer count, accuracy, and the share of fixes accepted without modification.
  • Documentation of how a running Foresight AI test is stopped and how an applied change is rolled back.
  • Early customer reports on whether Failure Atlas diagnoses hold on stacks unlike the ones in Gremlin's data.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence30
Adoption
Insufficient
Hype gap+45
Incentives80
Confidence35
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Gremlin launched Foresight AI on October 7th in a general-availability announcement.

    ReportedSupportedView cited source
  2. [2]

    The product identifies weaknesses, recommends or delivers configuration patches and infrastructure-as-code changes, then reruns the test that exposed the risk.

    ReportedSupportedView cited source
  3. [3]

    Gremlin says the product completed a beta, but the announcement provides no beta customer count, duration, incident-prevention results or other independent performance measures.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. runtimewire.com

    1 article · October 8, 2026

    Gremlin turns chaos tests into AI-generated fixes for reliability risks

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories