Skip to content

Invest1 publisher3 min readPublished

The UN dialogue's co-leads build their AI safety baseline from six existing frameworks

The co-leads of July's Geneva dialogue want models tested independently before release and incidents reported after it. Their list of things due by May is an evidence assessment, an agenda item and one cross-border sandbox pilot.

The Investor · Invest desk

Photograph accompanying The UN dialogue's co-leads build their AI safety baseline from six existing frameworks
Photo: yahoo.com

What happened

  • Bank of England Governor Andrew Bailey, chairing the Financial Stability Board, told G20 finance ministers and central bankers that today's already shaky global financial system faces a new threat in AI.
  • At the first UN Global Dialogue on AI Governance in Geneva in July, 170 countries came together with industry and civil society on mechanisms to begin governing the technology.
  • They set three deliverables for the May session: an evidence assessment from the scientific panel, an agenda that takes up the safety baseline, and a first cross-border sandbox pilot.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint A pre-release testing gate moves the release date out of the developer's hands, and the cost lands on anyone whose product plan assumes a model ships in a given quarter.
  • exposure The first binding questions about shared model and cloud dependence are likely to reach banks and insurers through their own supervisors, well before any UN text exists.
  • decision Someone has to capitalise the participation fund before May, or the baseline gets drafted by the delegations that can already afford the travel.
  • precedent Borrowing the Global Financial Innovation Network model would make regulator-to-regulator joint testing the default route for cross-border AI rules.

Independent pre-release testing puts someone other than the developer on the release schedule. "New AI systems should only be released after being independently tested against agreed thresholds of what they can do," the two co-leads wrote [4]. They also want developers to publish safety protocols and any post-deployment incidents monitored and reported [5], on the grounds that most AI research and development happens behind closed industry doors [12]. The op-ed does not say who would run the tests or how large the fund it asks for would be [22].

The dependence described in the piece is many institutions relying on the same handful of AI models and cloud providers, where one exploited weakness can spread rapidly [3]. The Geneva track's next fixed date is the May session [9]. That is eight months after the piece ran on September 21 [18], and ten months after the Geneva meeting itself [19].

The baseline is a compilation. It would be grounded in international human rights law and assembled from the OECD AI Principles, the G7 Hiroshima code of conduct, the Frontier AI Safety Commitments from the 2024 Seoul Summit, UNESCO's Recommendation on the Ethics of AI, and ISO and NIST standards [10]: six instruments, all already written [20]. The piece singles out the last two as the standards "engineers actually build to" [11]. The co-leads call the immediate steps "no regret" actions, steps they say carry no downside [16].

No item on the May list sets a date for pre-release testing [21]. Of the three items due in May, two are procedural: an assessment from the UN's Independent International Scientific Panel on AI, which presented its first one in Geneva [8], and a Dialogue agenda that takes up the safety baseline [9]. The third is one cross-border regulatory sandbox, piloted by a small group of countries and copied from the Global Financial Innovation Network [14].

Money, in the near term, goes to attendance. The co-leads want a dedicated trust fund on the IPCC model so governments, researchers and civil society with less financial capacity can be in the rooms where the conversations happen [13]. They offer Costa Rica's National AI Strategy, built with government, civil society, industry and academia at the same table, as the process to replicate [15].

I would expect binding language out of the financial-stability route before the Dialogue produces any, because Bailey's audience already supervises the institutions buying the models [2] while the Geneva audience next meets in May [9]. The baseline could be adopted quickly precisely because it restates commitments governments already made [10], in which case May yields text and no new obligation. Or the industry's own calls for a slower pace, now coming from several leading CEOs citing risks including loss of control, cyberattacks and bioterrorism [1], could harden into a dated testing rule. A May agenda with a date attached to independent testing would show the second.

What to watch

  • Whether the May Dialogue agenda puts a date on independent pre-release testing, or carries the safety baseline only as a discussion item.
  • Whether any government capitalises the IPCC-style participation trust fund, and at what size.
  • Whether the FSB follows Bailey's G20 remarks with third-party concentration work that names model and cloud providers.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories