Security1 distinct publisher2 min readPublished
Vedere Labs did get Claude Code to move a working exploit onto a second WAGO controller, and the price of doing it puts a measurable ceiling under the pitch that AI now writes OT exploits at machine speed.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
More than $500 in API usage across a session of more than eight hours works out to roughly $60 an hour in inference spend [13][15], and that covers only the final stage of the port. A researcher sat with it the whole way, supplying firmware context and pulling the model off incorrect leads [7].
No new vulnerability came out of this. CVE-2021-31886 is a pre-authentication buffer overflow in the Nucleus FTP server that lets an unauthenticated attacker run arbitrary ARM shellcode, and Forescout already had a working exploit for the WAGO 750-852 [2]. What was being measured is the porting labor: moving that exploit to the related 750-831 [3]. Confirmation was fast. Given a terminal, Ghidra, reference files and the physical device [5], Claude Code combined live probing with static firmware analysis and produced a payload that crashed the controller [6]. Crashing a PLC is the cheap part. Controlled execution stalled until the team moved from Claude Sonnet 4.6 to Claude Opus 4.6 and instructed the model to ask for help whenever it was unsure about a firmware detail [8], after which it worked out why its injected code kept being erased before it could run [9]. Two working payloads followed within 12 minutes [10], about 2.5 percent of the eight-hour session [16].
For asset owners the follow-up session carries more weight than the port. Asked to build a command-and-control implant, the model tested progressively more complex payloads until one wrote into a region mapped to the PLC's flash memory and permanently bricked it [12]. An agent that escalates writes against a live controller will eventually take the asset out, which costs an intruder the foothold and hands the operator the loudest indicator available.
Forescout says plainly that the guiding researcher could have done the port without AI in less time, at lower cost, and with the PLC still alive, and argues the question worth tracking is what happens as the required expert intervention falls, because AI can lower the marginal cost of repeating the work across many related targets at once [14]. That is the claim to hold them to. On this run the compression landed on the step that was already cheap: generating more payloads after a human had understood the memory behavior. The experiment ran in the wake of the coordinated attacks on water-sector PLCs [4]. Against that backdrop, the figure to track per firmware family is how many times a human had to intervene.
Ranked by verification strength, evidence, and original report placement.
Researchers at Forescout's Vedere Labs used Anthropic's Claude to port a working remote code execution exploit from one WAGO programmable logic controller to another, succeeding only after extensive researcher oversight, several hours of dedicated work, and hundreds of dollars in API costs.
The starting point was a previously developed exploit for the WAGO 750-852 PLC based on CVE-2021-31886, a pre-authentication buffer overflow in the Nucleus FTP server that allows an unauthenticated attacker to execute arbitrary ARM shellcode on the targeted PLC.
Forescout set out to adapt that exploit to a related but distinct model, the WAGO 750-831, and to see whether the AI could push the result further into a full command-and-control implant.
The experiment was conducted in the wake of the recent attacks targeting PLCs in the water sector.
The researchers used Claude Code, giving it access to a terminal, reference files, the reverse-engineering tool Ghidra, and the physical target device.
The AI confirmed the vulnerability through a mix of live probing and static firmware analysis before generating a payload that crashed the PLC.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Three Claude agents, one task, and a malware turf war: the multi-agent bill arrives1 distinct publisher
leadership
Same price, cheaper fast mode: Opus 4.8 argues on unit economics1 distinct publisher
build
A session that read "finished" and "still executing" was a slow queue, not a dropped handshake1 distinct publisher
build
Claude Code now opens in auto mode: a classifier, not you, approves the shell commands1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one lab's own account
Every figure that matters here — the $500, the eight hours, the twelve minutes, the bricked controller — comes from Forescout describing its own experiment, relayed by SecurityWeek. The underlying flaw is a public CVE and the method is described in enough detail to be argued with, which is more than most vendor research offers. What is missing is anyone outside Vedere Labs who has re-run the port or checked the meter.
A single bench run
What exists is one experiment, one controller model, one dead device at the end of it. Nothing in this reporting shows the technique used in a real intrusion, reproduced by a second team, or folded into anyone's tooling — and the water-sector PLC attacks that set the timing are not attributed to AI-assisted exploit work.
Priced below the pitch
This reporting cuts against the machine-speed-exploit story rather than riding it: it says outright that the human would have been faster, cheaper and less destructive, and it puts a dollar figure on the alternative. The one place it reaches is forward-looking — the premise that expert babysitting keeps getting cheaper — and even that arrives phrased as a question rather than a finding.
Vendor lab, unflattering result
Vedere Labs is the research arm of a company that sells OT security, and an AI-writes-PLC-exploits headline lands well for that business — the framing alongside fresh water-sector PLC attacks is a choice, not a coincidence. What complicates the read is the shape of the result: hundreds of dollars, a wrecked controller, and an admission that a human would have done it better. Marketing rarely leads with that.
Concrete but unreplicated
Specific enough to believe about what happened on that bench; thin enough that we would not build a threat model on it. One publisher, one lab's telling, one device, and a cost figure that covers only the final stage of the work — with the earlier stages, and any second opinion, absent.