Skip to content

Science1 publisher3 min readPublished

OpenAI spells out the access it gives outside assessors of its safety cases

The company says past engagements have included visible chain of thought and confidential data, and its new post proposes the four areas it wants outside assessors to examine and the three questions they should ask.

The Scientist · Science desk

Illustration accompanying OpenAI spells out the access it gives outside assessors of its safety cases

What happened

  • OpenAI has published a set of priorities and principles for third party safety assessments, proposing four priority areas for deeper examination along with principles for independent, rigorous and secure work.
  • The engagements it describes run long, with some lasting weeks and others several months, and the company expects to support several of them in parallel over different periods of time.
  • The scope is private and non-profit assessment organisations doing technical safety work, held separate from OpenAI's testing and evaluation work with governments.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint A launch-agnostic assessment is not a release gate. An assessor who finds a problem mid-study has nothing scheduled to hold while the company decides what to ship.
  • decision Assessment organisations have to decide whether access on the subject's terms is worth taking when the subject also proposes the priority areas and the questions worth asking.
  • precedent A published list of what access a lab will grant gives regulators and other labs a concrete reference point for what an AI audit is supposed to include.

Visible chain of thought, together with internal deployment access for incident response and monitor red teaming, lets an outside team check whether a monitor flags the behaviour it is claimed to flag [2]. No volume of API querying produces that view. OpenAI says the access should let assessors challenge its assumptions, identify risks it may have missed, and reach their own conclusions about how well the safeguards work [3].

OpenAI sets the questions as well as the access. It says assessments are most useful when they address specific, consequential questions, and names three: whether the evidence supports a lab's safety case and claims, whether evaluations adequately test the risks they are intended to measure, and whether safeguards work under realistic conditions [4][5]. Those are the right questions to ask of a safeguard, and they are also the subject's framing of what counts as a finding.

OpenAI writes that labs have a responsibility to enable meaningful scrutiny while protecting sensitive information, and that labs and independent assessors share the responsibility for getting this right [10]. The post does not say who decides where that boundary falls, or which organisations have held the access so far [13].

The definitions are the testable part: a safety claim, as the post defines it, should identify the risks and conditions it addresses along with its assumptions and limitations [6]. A safety case connects individual claims to evidence and makes explicit the assumptions, uncertainties and remaining risks that could affect its conclusions [7].

What these engagements can do depends on their timing. Some run weeks and others several months, multiple in parallel, and the work is described as launch-agnostic, aimed at particular safety claims in depth over time [8][9]. The post allows that third party assessments can also be part of pre-deployment work and may inform deployment decisions [8]. That is a separate track from the one it sets out here. A finding from a launch-agnostic study arrives when the study ends, and the deployment it bears on may have shipped already.

The scope is private and non-profit assessment organisations doing technical safety work. OpenAI says its work with governments on testing and evaluation may call for different approaches because the roles and responsibilities differ [11]. Third party assessments are already part of the company's Preparedness Framework practices, and it says it has supported organisations and legislation pushing for a more rigorous and accountable process [14]. The first of the four priority areas is independent assessment of safety cases across training, evaluation, internal deployment and external deployment. It asks for expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse, and red teaming, with multiple assessors likely examining different parts [12]. OpenAI says labs and assessors should be operating with shared international standards for safety and security practices [10].

What to watch

  • Whether a named assessment organisation publishes a review of an OpenAI safety case, and whether OpenAI publishes findings it disagrees with.
  • Whether any of these long-running assessments is ever tied to a release decision instead of running launch-agnostic.
  • Whether a standards body names the shared international standards for safety and security practices the post invokes.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories