Skip to content

Science2 publishers2 min readPublished

Google, OpenAI and Anthropic reportedly plan their own body to set frontier AI's pre-release tests

Google, OpenAI and Anthropic are reportedly building SAFA, a government-independent body to set pre-release AI testing rules, aiming to launch in early 2027. Its members, evaluators and enforcement powers are unannounced, so buyers still have to set testing terms with each vendor themselves.

The Scientist · Science desk

Illustration accompanying Google, OpenAI and Anthropic reportedly plan their own body to set frontier AI's pre-release tests

What happened

  • The Information reports that Google, OpenAI and Anthropic are working together on a body tentatively named the Standards Authority for Frontier AI, or SAFA.
  • As described, SAFA would operate outside government control and write guidelines for risk assessment, testing and pre-release review of frontier models.
  • People close to the plan say the aim is to launch the body officially in early 2027.
  • The report came in the same week that Anthropic's Dario Amodei and OpenAI's Sam Altman urged the United Nations to create safeguards around AI.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint Until someone is named to run the evaluations and a consequence is set for failing them, a SAFA result cannot replace a buyer's own due diligence on a model.
  • decision Buyers who want pre-release assurance before SAFA exists have to write it into their vendor contracts themselves, because the body would arrive in late 2026 at the earliest.
  • precedent If SAFA launches as described, the first shared pre-release baseline for frontier models would be written by the developers it covers, so the charter's terms on who can join and who evaluates will decide how much weight it carries.
  • contradiction OpenAI also advocates mandatory US safety rules, so the idea that industry would set the rules in place of regulators matches only half of one founder's stated position.

A testing standard is only useful to the people downstream of it if two questions are settled: who runs the test, and what a failing result triggers. According to Superpower Daily, the companies have not announced membership, standards or enforcement powers, and the proposal assigns neither the evaluator's job nor the consequence of falling short [6][7]. A shared method would let buyers compare how models were tested, while leaving open whether anyone outside the developer checked the result [6]. Superpower Daily also gives a wider launch window than CIO, late 2026 or 2027 [4].

An industry venue already exists. All three companies belong to the Frontier Model Forum, founded in 2023 to advance safety research and best practices [5]. Superpower Daily asks whether SAFA would develop testing methods, assess whether members meet them, or simply publish guidance, and notes that those are different jobs [5]. A forum for swapping research need not hold authority over its members, the outlet adds. A standards body might do more, but its name cannot by itself establish what it would require of a developer whose model falls short [8].

OpenAI's blog post this week put industry standards next to government bodies. The company called for agreed-upon baselines, common measurements and incident reporting protocols, and wrote that such international standards "may be as important to pacing the frontier as alignment research itself" [9]. It named the US Center for AI Standards and Innovation, other federal frameworks, state regulations and new public-private partnerships as groups that can work together on that goal [10].

In the account of independent analyst Carmi Levy, buyers care less about global AI safety than about day-to-day exposure: agents leaking corporate or employee data, platform vulnerabilities, hallucinations in workflows, and compliance [12]. "They're worried about who's responsible for what, when (not if) AI goes off the rails," Levy said [13]. After a summer of reports of agents breaking containment, he said, AI companies are in no position to give that assurance [14]. He advised enforcing vendor standards now. One option is to require every new model to ship with the "equivalent of a safety and security datasheet" setting out its capabilities, known failure modes and testing history [15].

I think the reported design supports half of the idea that industry would set the rules buyers come to lean on. If SAFA launches as described, the labs would write the pre-release baseline for their own models, outside government control [2]. Whether buyers can rely on it depends on choices Superpower Daily lists as still open: who joins, how smaller developers are treated, and how independent any evaluators are [16].

What to watch

  • Whether SAFA's charter names who runs pre-release evaluations and requires members to act when a model fails them.
  • Whether SAFA's remit is formally separated from the Frontier Model Forum's, and whether smaller developers can join on equal terms.
  • Whether US lawmakers take up the mandatory safety requirements OpenAI advocates, or CAISI gets a role in SAFA's testing.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories