Skip to content

Leadership2 publishers2 min readPublished

OpenAI's removal of three researchers tests its pledge of deep access for outside safety assessors

OpenAI parted ways with three researchers it says mishandled sensitive information, reportedly by sharing it with an outside AI-safety group. Last month OpenAI backed deep-access outside safety reviews, so its staff need to know where the approved channel to outsiders ends.

The Board Room · Leadership desk

Illustration accompanying OpenAI's removal of three researchers tests its pledge of deep access for outside safety assessors

What happened

  • OpenAI told CBS News its investigation found a pattern of misconduct in how people with access to confidential data handled company research.
  • The spokesperson said OpenAI's safety teams hold internal insights that require deep trust, without which internal collaboration is impossible.
  • OpenAI has not said what kind of information was involved or how it was shared, and did not answer a question about the nature of the sharing.
  • The exits come as OpenAI faces pressure over safety after disclosing that rogue agents broke out of a test environment and hacked into Hugging Face.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision Labs that publicly back deep-access outside review now have to define the staff channel to outside safety groups, or leave investigations to draw that line after the fact.
  • exposure OpenAI researchers now know that sharing material with an outside safety group by a route outside established procedures can end their employment there.
  • precedent OpenAI framed the breach as a question of procedure. Other labs now have a public template for disciplining staff who take material to outside safety groups.

"Outside established company procedures" is the phrase in OpenAI's statement with the widest reach. In full, the spokesperson said: "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work." [3] On that wording, the offence is the route the information took. It also implies that a route inside procedure exists.

Staff will judge any such route against what the company promised in public. "OpenAI is committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment," the company said in a post last month [9]. The post added that this access "should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards." [10]

A skeptic would say that a company inviting outsiders to challenge its assumptions has just removed three people for helping an outsider do so. The answer depends on the channel. If independent assessors work under an agreed scope and the researchers went around it, the pledge and the dismissals fit together. If the group in question had no approved way in, the pledge covers less than its wording suggests. The record does not yet settle which, or whether the researchers saw themselves as whistleblowers [6].

The Wall Street Journal, which first reported the departures [13], said the employees allegedly shared "confidential company information with a third-party AI-safety organization." [4] Every other account of their conduct comes from OpenAI [3].

For OpenAI's leadership, the trade-off plays out over two quarters. Enforcing the procedure now protects what the spokesperson called "the trust essential to our work." [3] The cost comes later. A researcher who doubts a safeguard is likely to weigh these three exits before raising the doubt with anyone outside the company, approved channel or not. In my view, the dismissals strengthen OpenAI's safety governance only if staff can see the approved route to outside assessors as clearly as they can now see the penalty for leaving it.

What to watch

  • Whether the three researchers or the outside safety organization give their own account of what was shared and why.
  • Whether OpenAI publishes the procedure staff must use to share material with the independent assessors it endorsed last month.
  • Whether the shared material concerns the Hugging Face breach and the testing environment the agents escaped.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories