Build1 publisherNot yet confirmed elsewhere2 min readPublished
Gremlin's Foresight AI reruns the failing chaos test to check its own proposed fixes
Gremlin released Foresight AI on October 7th to propose fixes for the faults its chaos tests expose. The launch came with no beta accuracy or acceptance figures, so reliability teams have only their own supervised trials to judge whether those fixes hold.
The Engineer · Build desk

What happened
- Gremlin's homepage says Foresight AI proposes tests and changes for the team to approve, keeping a human sign-off step in the loop.
- Diagnoses draw on the test, service and health-check context, with each recommendation tied to observed test results, according to Gremlin's documentation.
- A proprietary Failure Atlas of more than a decade of cause-and-effect data on failures and recovery underpins the recommendations, Gremlin says.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Reliability teams can adopt Foresight AI under the approval gate and use their own accept, edit and reject counts as the first accuracy evidence they can check.
- exposure Granting the external-model permission lets service context reach an outside AI service, so the data-flow review has to clear before anyone judges diagnosis quality.
- cost Buyers carry the work of confirming how to halt a running test and revert an applied change before any proposed patch reaches a production system.
Foresight AI starts from a chaos experiment that has already exposed a weakness. Gremlin positions it against AI tools that respond to incidents already underway, aiming instead at failure modes a team can test in advance [10]. Going by Gremlin's materials, a run has four steps:
1. A controlled fault test exposes a weakness in a service [2]. 2. The product diagnoses it from the test, service and health-check context, tying its recommendations to the observed results [6]. 3. It proposes a configuration patch or an infrastructure-as-code change for the team to approve [2][8]. 4. It reruns the test that exposed the risk, and can repeat tests as the system evolves [2][11].
Step four is the good engineering here. Runtimewire put the case plainly: a recommendation can be plausible and still leave the failure mode untouched, and rerunning the originating test checks the result under the same condition [13]. A passing rerun covers one injected fault. Side effects of the change on paths that test never exercised have to surface in the rest of the suite.
Step two draws on what Gremlin calls its Failure Atlas, a proprietary set of more than a decade of cause-and-effect data on failures and recovery [5]. Treat it like any benchmark built on someone else's workload. For its diagnoses to transfer, the failures in that data have to resemble the dependencies and recovery behaviour of the stack under test.
The model question is separate. Use of an external large language model is optional and requires customer permission, Gremlin says [7]. Runtimewire notes that buyers are able to check which system context the product draws on and under what conditions it may be passed to an AI service [14]. Granting that permission is a different decision from approving any single change.
Gremlin says Foresight AI completed a beta. The announcement also leaves out the beta's customer count, its duration and any incident-prevention results [3][9]. Runtimewire concluded that Gremlin's claim to have already proven the product can identify and address risks before incidents is unverified by those measures [4].
I think the right deployment is a supervised pilot on non-critical services. Keep the approval gate Gremlin describes on every change [8], and log each outcome as accepted, edited or rejected. Counted over the pilot, those outcomes give an acceptance rate measured on the team's own services.
What to watch
- Gremlin publishing beta figures: customer count, accuracy, and the share of fixes accepted without modification.
- Documentation of how a running Foresight AI test is stopped and how an applied change is rolled back.
- Early customer reports on whether Failure Atlas diagnoses hold on stacks unlike the ones in Gremlin's data.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives80
- Confidence35
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Gremlin launched Foresight AI on October 7th in a general-availability announcement.
- [2]
The product identifies weaknesses, recommends or delivers configuration patches and infrastructure-as-code changes, then reruns the test that exposed the risk.
- [3]
Gremlin says the product completed a beta, but the announcement provides no beta customer count, duration, incident-prevention results or other independent performance measures.
- [4]
Gremlin's claim that Foresight AI has already proven it can identify and address risks before incidents remains unverified by those measures.
- [5]
Gremlin says its proprietary Failure Atlas draws on more than a decade of cause-and-effect data about system failures and recovery.
- [6]
Gremlin's Reliability Intelligence documentation describes diagnoses based on the test, service and health-check context, with recommendations tied to observed test results.
- [7]
Gremlin says use of an external large language model is optional and requires customer permission.
- [8]
Gremlin's homepage says Foresight AI proposes tests and changes for the team to approve.
- [9]
The announcement does not quantify the beta's accuracy or say how often suggested fixes were accepted without modification.
- [10]
In an August 6th explanation, Gremlin contrasts AI tools that respond to an incident already underway with Foresight AI's focus on failure modes teams can test in advance.
- [12]
Buyers running mission-critical systems will still need evidence of accuracy and a clear account of safety controls, including how to stop a test or roll back a change.
- [13]
A recommendation can be plausible and still leave the underlying failure mode untouched; rerunning the originating test gives engineers a way to check the result under the same condition.
- [14]
Infrastructure buyers can assess what system context the product uses and when that context can be shared with an AI service.
Sources
1 independent publisher whose own reporting we read for this story.
- runtimewire.comGremlin turns chaos tests into AI-generated fixes for reliability risks
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Chaos EngineeringFollow
- AI-assisted remediationFollow
- Site reliability engineeringFollow