BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Gremlin's Foresight AI agents wait for an engineer before every test and fix
Gremlin has made Foresight AI generally available, an add-on whose agents plan chaos tests, run them, propose fixes and rerun the original test. Every test and every fix still needs an engineer's sign-off, so human review moves to an approval queue.
The Engineer · Build desk

Gremlin says Foresight AI will neither run a test nor apply a remediation without approval. Wix has described AI interpreting logs and suggesting remediations, with deployment kept under human control. Cloudflare's proposal envisages agents managing more of testing, deployment and maintenance.
- decision Gremlin Foresight AI Gremlin says it will neither run a test nor apply a remediation without approval, claim 7
- precedent Wix Has described AI interpreting CI/CD logs and suggesting remediations, with deployment decisions kept under human control, claim 10
- precedent Cloudflare Agent Development Lifecycle proposal envisages agents managing more of testing, deployment and maintenance, claim 11
| Who | How | Kind | Claim |
|---|---|---|---|
| Gremlin Foresight AI | Gremlin says it will neither run a test nor apply a remediation without approval | decision | 7 |
| Wix | Has described AI interpreting CI/CD logs and suggesting remediations, with deployment decisions kept under human control | precedent | 10 |
| Cloudflare | Agent Development Lifecycle proposal envisages agents managing more of testing, deployment and maintenance | precedent | 11 |
What happened
- Recommendations draw on Gremlin's proprietary Failure Atlas, which the company says spans millions of fault-injection experiments across tens of thousands of systems.
- Gremlin says system data does not train a language model, and the LLM is limited to searching, summarising and explaining results.
- A fourth agent, styled a Technical Program Manager, tracks coverage, risks and commitments across teams and sends weekly reports through Slack.
Why it matters
- constraint Per-item approval caps throughput at the rate engineers can sign off, the same human review capacity Gremlin says AI-assisted delivery is already outrunning.
- exposure Recommendation quality rests on how closely a buyer's failure modes match a corpus outsiders cannot inspect, so only a trial on real services tests the fit.
- decision Teams adopting agent-driven testing have to pick where the human boundary sits: before each test as Gremlin does, at deployment as Wix does, or further out as Cloudflare proposes.
The part worth copying is the loop. When a test exposes a weakness, the Operator agent proposes a change and the service reruns the original test [5][6]. The proof that a fix worked is the experiment that failed, now passing. Andrus said: "An alert going quiet doesn't prove anything. A test that passes under load, in realistic conditions, does." [8] We think rerunning the failed test is the right standard. A quiet alert shows a symptom stopped. A passing fault injection shows the failure mode is handled under the conditions that test models.
The fourth agent, a Technical Program Manager, tracks test coverage, reliability risks and commitments across teams, then writes the weekly report and posts Slack updates [5]. Few engineers will miss doing that job. The service also builds dashboards from natural-language requests, meant to show changes in reliability scores and the risks teams have addressed [9].
The second sound decision is where the language model sits. According to Gremlin, the agents decide from platform data and the Failure Atlas, and the LLM handles search, summaries and explanations [16]. That keeps a text generator out of the step that chooses which fault to inject. Gremlin also says system data is not used to train a model [16].
Treat the Failure Atlas like a benchmark table. Gremlin describes it as a proprietary record of millions of fault-injection experiments run over more than ten years across tens of thousands of distributed systems [15]. Its recommendations transfer to a service only if that service fails the way the corpus did, with comparable dependency graphs and comparable retry and timeout behaviour. Because the record is proprietary, a buyer can only test the fit by running it against their own services [15].
Gremlin says Foresight AI will neither run a test nor apply a remediation without approval [7]. It positions the product against AI-assisted delivery, with Andrus arguing that more code is reaching production without close human review [4]. In InfoQ's account he does not attach a figure to that [4]. So a product aimed at a review shortfall still asks a person to approve every test and every fix. The engineer moves from writing experiments to reviewing them. We'd expect a review to cost less than authoring one, but the approval queue still grows with each service enrolled.
Wix has described using AI in CI/CD for log interpretation and remediation suggestions, with deployment and critical infrastructure decisions left to humans [10]. Gremlin's gate covers running the test as well as applying the fix, and we think it belongs there for a first GA release of a tool whose job is injecting failures [7]. Cloudflare's Agent Development Lifecycle proposal has agents managing more of testing, deployment and maintenance [11].
Gremlin's own claim about expertise is narrow. The product is meant to provide reliability checks without every developer becoming a distributed-systems specialist [3]. Engineer Sandip Bhattacharya argued in a LinkedIn discussion that organisations still need specialists to configure agents [12]. Intellyx analyst Jason English, in an article written for Gremlin, cautioned that a reliability programme cannot simply be installed as a software package and needs shared ownership and accountability across engineering and management [13]. Intellyx discloses that Gremlin is its customer [14].
What to watch
- Whether Gremlin adds a mode that lets agents run tests or apply fixes without per-item approval, moving toward the autonomy in Cloudflare's lifecycle proposal.
- Customer reports, from sources Gremlin does not pay, on how often Operator-proposed fixes pass the rerun test.
- Whether Gremlin opens any part of the Failure Atlas to outside inspection.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Gremlin announced the general availability of Foresight AI, an agentic product that analyses services for potential failures, recommends changes and reruns tests to check that a fix works.
- [2]
Foresight AI is available as an add-on to the Gremlin platform.
- [3]
Foresight AI is intended to provide reliability checks without requiring every developer to become a distributed systems specialist.
- [4]
Gremlin positions the product as a response to faster AI-assisted software delivery, with CEO and co-founder Kolton Andrus arguing that more code is reaching production without close human review.
- [5]
Foresight AI divides work between four agent roles: an Analyst gathers service data and recommends tests; a Tester schedules and runs them; an Operator interprets failed tests and proposes remediations; a Technical Program Manager tracks test coverage, reliability risks and commitments across teams, produces weekly reports and sends updates through Slack.
- [6]
When a test uncovers a weakness, the service proposes a change and then runs the original test again (a closed validation loop).
- [7]
Gremlin says Foresight AI will neither run a test nor apply a remediation without approval.
- [8]
"An alert going quiet doesn't prove anything. A test that passes under load, in realistic conditions, does."
- [9]
The service creates reports and dashboards from natural-language requests, including charts, tables, tabs and links, intended to show changes in reliability scores and the risks teams have addressed.
- [10]
Wix has described using AI in CI/CD for log interpretation and remediation suggestions while keeping deployment and critical infrastructure decisions under human control.
- [11]
Cloudflare's Agent Development Lifecycle proposal envisages agents managing more of testing, deployment and maintenance.
- [12]
In a LinkedIn discussion about AI-based SRE automation, engineer Sandip Bhattacharya argued that organisations still need specialists to configure agents.
- [13]
In an article written for Gremlin, Intellyx analyst Jason English cautioned that a reliability programme cannot simply be installed as a software package and that sustained results need shared ownership and accountability across engineering and management.
- [15]
Foresight AI bases its recommendations on Gremlin's Failure Atlas, which the company describes as a proprietary record built from millions of fault-injection experiments over more than ten years across tens of thousands of distributed systems.
- [16]
Gremlin says system data is not used to train a large language model; the agents use platform data and the Failure Atlas for decisions, with an LLM used for searching, summarising and explaining results.
Sources
1 independent publisher whose own reporting we read for this story.
- infoq.comForesight AI Brings Gremlin Agents to Reliability Engineering
1 article · October 11, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI agents in software deliveryFollow
- Chaos EngineeringFollow
- Site reliability engineeringFollow