Leadership1 publisher2 min readPublished
OpenAI's CFO described a training saving that Silicon Valley heard as self-improving AI
Sarah Friar told a Goldman Sachs conference that OpenAI's biggest models now train its smaller ones to hold down training costs. In the same week, two Anthropic safety leads put their reasoning about extinction risk in writing.
The Board Room · Leadership desk

What happened
- OpenAI CFO Sarah Friar told a packed Goldman Sachs tech conference in San Francisco on Tuesday that the company's biggest models can now train smaller ones, saving on the cost of training them.
- Analysts and investors in the room heard a business opportunity, while some AI researchers treat the same capability as terrifying, according to Business Insider.
- Employees of Anthropic and OpenAI posted about the risk of AI killing humans this week, setting off what Business Insider describes as a firestorm across the valley.
- Business Insider reports Wall Street puzzling over why Anthropic staff would say their main product will kill humans right before what could be the biggest IPO in history.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- contradiction A cheaper small model and a more capable successor are different outcomes, and the same word was applied to both, so anyone pricing the disclosure has to work out which one OpenAI actually described.
- cost The saving Friar described lands on OpenAI's own training bill, and a customer reading it as a forward signal on token prices is reading in a promise about token prices that Friar's remark did not make.
- decision Leaders who have kept out of the catastrophic-risk argument now have named employees of their vendors publishing numbers and motives. Saying nothing in a vendor review is itself a position.
- exposure Whoever drafts Anthropic's risk factors ahead of a listing has to reconcile them with what the company's own safety leads post under their names.
Friar's remark and the label attached to it describe two different things. Business Insider defined recursive self-improvement as AI getting powerful enough to create the next, even more powerful system itself, and treated bigger-models-training-smaller-models as an example of it [3]. Training a smaller model from a larger one lowers the cost of the smaller one. Building a stronger successor raises capability. Friar is not quoted using the term herself [1].
The incentive account came from inside Anthropic, in writing. Samuel Marks, scalable oversight lead at the company, wrote on X: "Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely" [6]. Evan Hubinger, Anthropic's alignment science lead, wrote: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade" [5].
Hubinger's figure covers ten years. That span matters to anyone converting it into a plan. Spread evenly, a cumulative probability of a little over 10% across a decade works out to an annual rate of roughly 1.05 percent: 1 - (1 - 0.0105)^10 = 0.10 [12].
Brad Gerstner of Altimeter Capital, an investor in both Anthropic and OpenAI, said "It's fucking nonsense" when asked whether AI will kill humans [7]. He is answering the extinction claim. The other item in the same account can be tested: agents breaking out of their sandbox testing environments [8]. A company running agents can look for that in its own logs.
Kylan Gibbs, CEO of Inworld, told Business Insider that the week's extreme doomerism is driven by a growing loss of agency, with control over development concentrating in a few labs and researchers inside them feeling little influence over the direction [9]. Sam Altman said on a podcast earlier this year: "People want agency, self-determination, the ability to play a role in architecting the future alongside the rest of society" [10]. The agency a buyer has this quarter runs to containment testing and contract terms. Hubinger's decade estimate is one researcher's number, posted on X [5].
What to watch
- Whether OpenAI publishes which direction its model-trains-model pipeline runs, in a technical report or a filing.
- Whether Anthropic's IPO paperwork, if it comes, includes the risk language its own safety leads publish on X.
- Whether more lab employees or executives put numeric odds on the record after this week's exchange.