Build1 publisherNot yet confirmed elsewhere3 min readPublished
OpenAI ships new capability and risk every Tuesday, its former safety-report lead says
David Robinson, who led OpenAI's model safety reports, says reasoning training and new tools now change its models weekly with no new pretraining run. On his account, teams building on OpenAI cannot assume a model name fixes what the system does or what risks it carries.
The Engineer · Build desk

What happened
- GPT-6.1 Sol arrived at DevDay a week after GPT-6 Sol, the account's example of how short the gaps between OpenAI releases have become.
- Robinson said system cards made sense when frontier models arrived every few months, and argued for a live dashboard tracking safety properties from predeployment testing through post-release behavior.
- Robinson announced his resignation in an Oct. 3 essay for The Atlantic, and Sam Altman replied on X that OpenAI is working to keep its models from outpacing its safety measures.
- Robinson said OpenAI research teams now use more than 100 times the agentic compute they used at the start of the year, calling the effect of coding agents "night and day".
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Builders need a re-test trigger other than a new model name, because Robinson places capability changes in reasoning training and tools as well as in the base model.
- exposure Applications built on OpenAI models can pick up new behavior and new risk from a tool integration even when no training run sits beneath the change.
- constraint A system card describes one snapshot, so a builder relying on it may be reading about a system that has since gained tools or a new reasoning recipe.
When Robinson started at OpenAI in May 2023, a day after Sam Altman's first Senate testimony, shipping a frontier model generally meant training one from scratch [1]. "We were going to bake a fresh cake with a new pretraining run, do the whole thing from scratch," Robinson told Ezra Klein on "The Ezra Klein Show" [2]. That process took months and held major releases to a few a year [3]. Post-training and safety testing came only after each new base was finished [3].
His description of the current stack separates the layers by cost. "You've got the pretraining; that's the baking of the underlying model," Robinson said [8]. He said "But then, in addition to post-training, you have reasoning training" and that "those steps are easier to do quickly, so you can redo them" [9]. This is sound engineering: if the bottom layer takes months, you design the upper layers so they can be replaced without touching it. A better reasoning recipe can be layered onto the base model OpenAI already has [6].
The second layer is tooling. "It's not just a chat anymore," Robinson said of the tools and affordances now wired into models [11]. Those can make a system more capable and change how it behaves with no new training run beneath it [11]. "All of those things are changing what the model can do and what the risks are, and we're shipping new capability and risk every Tuesday," he said [12].
"Every Tuesday" is Robinson's characterization of the pace. OpenAI once made a point of the "month or months" of safety work that went into GPT-4 after the model was finished [7]. That window is longer than the whole gap between GPT-6 Sol and GPT-6.1 Sol [18]. It does not prove GPT-6.1 Sol got less testing, because testing for one release can run while the previous one ships.
For a team building on these models, the base model is now only one of the layers that can change between major releases [6]. The interview does not say which of those changes reach a model a developer calls through the API. In my context, where an application's behavior is the product, I would treat each OpenAI release as a dependency upgrade and run our own evaluations against it before users see it.
The third accelerator is OpenAI's own research speed. Its September report said that by mid-August the organization was running 3.1 eight-hour agent workdays for every human workday [15]. That is about 24.8 agent-hours for each human workday [19]. The least glamorous evidence is the easiest to believe. Robinson said the research infrastructure can be "pretty janky" [16]. Researchers who once took a broken experiment to an internal Slack channel now ask Codex, and OpenAI's report shows traffic to that channel falling as agent use climbed [16].
Competition adds pressure to use that speed. Klein pointed to a more crowded field than OpenAI faced a few years ago, with Anthropic, xAI, Chinese labs and open-weight models all competing for enterprise customers [17]. "All of these things push toward speed," Klein said [17].
What to watch
- Whether OpenAI builds the live safety dashboard Robinson proposed or keeps publishing one system card per model.
- Whether OpenAI release notes start saying which layer changed in a release: the base model, the reasoning training, or the tools.
- Whether OpenAI's response to Robinson's Atlantic essay goes beyond Altman's post on X.