Skip to content

Product1 publisher3 min readPublished

An OpenAI model found an 8.8 token-replay flaw in code for an EU project

Politico reported that ENISA and CERT-EU ran an advanced OpenAI model over an EU project's code and got four fixed flaws out of it. Poland's CERT, doing the same kind of work, said it tested every hypothesis on real systems.

The Product Desk · Product desk

Illustration accompanying An OpenAI model found an 8.8 token-replay flaw in code for an EU project

What happened

  • ENISA and CERT-EU ran a security analysis of an EU project's code with an advanced OpenAI model and found four flaws, spokesperson Laura Heuvinck told Politico, and the flaws have since been fixed.
  • The high-risk finding is CVE-2026-73431, scored 8.8, in Vulnerability-Lookup up to 5.5.1, where activation and recovery tokens were never marked as used and could be replayed while still valid.
  • The European Commission said on 10 September that ENISA is now testing Mythos 5 and GPT-6 Astra, after the EU spent months asking for access to the most capable US models.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Anyone proposing AI code review internally now has a named public-sector case with a CVE number and a described flaw behind it, so the argument moves off vendor benchmarks and onto whether your own reviewers can verify what the model raises.
  • constraint CERT Polska's caveat limits how much can get through: the models produce candidates and staff confirm them on live systems, so reviewer hours decide how much output ever becomes a fix.
  • exposure Manufacturers now file 24-hour early warnings through the EU's new reporting platform, and missing those duties reaches the act's top fine band. The platform's own code quality sits inside their compliance path.
  • precedent The four flaws came out of a scanning programme that has been running against EU institutions since July, so more EU advisories carrying model provenance are the expected output.

The 8.8 is a missing state check. According to the OpenCVE record Politico tied to the finding, the software never registered that an activation or recovery link had been used. An attacker holding one live link could replay it while it was valid and reset the password again [5][6]. The affected code is Vulnerability-Lookup, an open-source tool for tracking vulnerabilities, in versions up to 5.5.1 [4]. There is no crash in it and no memory corruption. The failure is a rule the code never enforced. TNW, which summarised the Politico report, said it had not independently verified it [8].

CERT Polska published six MikroTik RouterOS flaws on 5 September. The team said the models sped up the analysis but that it still tested every hypothesis on real systems, and that its own staff judged the impact of each flaw [9][11]. Two of those six are already being chained by attackers in a combination the team calls MikroTrick, which gives full control of routers that expose SSH to the internet. CERT Polska has seen the attacks since at least 2 September [12][13]. The advisory came three days later [14].

Across the two teams the published count is ten flaws found with an OpenAI model, four from ENISA and CERT-EU and six from CERT Polska [15]. Two of the ten are under active attack [12][15].

OpenAI's head of policy for Europe, Tom Duff Gordon, spoke of a "narrowing window" for AI to find weaknesses before attackers do, in a statement quoted by Politico [16]. "That's why we work with partners such as ENISA," he said [17]. ENISA is in OpenAI's trusted access programme for cyber models, which the company runs through its Daybreak cyber defence programme [18]. Heuvinck told Politico that ENISA and CERT-EU have used advanced models since July to "actively scan" EU institutions, and under a Commission action plan published the same month the agency and the Joint Research Centre are building a "secure testing platform" for the most advanced models [31][32].

The second review is AISLE's. AISLE said on 14 September that it had reviewed the code of the Single Reporting Platform for the Cyber Resilience Act and will keep reviewing it for 12 months [22][23]. Manufacturers of products with digital elements sold in the EU must now use that platform to report exploited flaws and severe incidents [25]. The Register reported the deadlines: an early warning within 24 hours, a fuller notice within 72, and a final report within 14 days of a fix, with the 24-hour deadline applying wherever the manufacturer is based [26]. Breaching those duties can bring the act's top fines of 15m euros or 2.5% of annual turnover, whichever is higher. Most of the act's other rules apply from 11 December 2027 [27][28]. ENISA's Hans de Vries said the agency had also run user and security testing with stakeholders including national CSIRTs. He added: "I thank AISLE for their important AI-based secure code review performed" [29][30]. The report does not say what the AISLE review found.

For whoever has to defend this to a security lead, the question is what has to be true before a model-flagged finding leaves the tracker: reproduction on a real system, and a named person who signs the severity. CERT Polska published both of those alongside its model names [11]. ENISA's four are fixed, and the 8.8 has a record dated 12 August behind it [2][7].

What to watch

  • Whether the secure testing platform ENISA is building with the Joint Research Centre publishes evaluation results or only grants model access.
  • Whether the AISLE review of the reporting platform produces published findings before its 12 months are up.
  • Whether EU authorities get the newest Mythos version. Politico reported they still lack it.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories