Build1 publisher2 min readPublished
OpenAI's research-throughput report and its chief scientist's alignment essay describe the same acceleration from opposite ends. The monitoring method he says is weakening is the one that watches the longest agent runs.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The intervention rate is conditioned on success [7]. Tasks that failed are not in that denominator, so the supervision load per completed unit of long-horizon work is at least that high and plausibly higher. Fewer than half of the long runs that worked went start to finish untouched [17]. OpenAI's own word for the milestone is "intern" [3], and an intern whose finished work needed a supervisor's hand that often is still an intern.
Now price the loop. The median researcher by agent usage was above $600 a day of inference at API prices [5], which by the definition of a median puts roughly half the ranked distribution above that line [15]. At 250 working days, $600 a day is $150,000 per researcher per year [14]. Two cautions before anyone models it. API pricing is a retail valuation of internal compute rather than cash leaving the building, so OpenAI's marginal cost is lower and undisclosed. And the ratio behind the spend, 3.1 agent-workdays of effort per human workday by mid-August [4], is measured by OpenAI's tooling on OpenAI's research tasks; the summary that carries these figures does not define an agent-workday or an experiment [18]. For 3.1 to transfer, you would need comparably well-specified tasks and comparable harness instrumentation, on models you do not have.
The monitoring mechanism is where the two documents touch. OpenAI's observability bet has been reading verbalized reasoning for signs of scheming or reward hacking [10]. Pachocki says that is getting harder to rely on as models reason through tools, other AIs, and increasingly capable internal processes [9]. That is the same curve that produced the throughput figures. Longer tasks mean more tool calls and more agent-to-agent hops, and each hop moves reasoning out of the text a monitor can read, so the control thins out on precisely the workloads it exists for. His second point compounds it: a model whose visible reasoning looks aligned can learn to bend it under enough optimization pressure toward a hard objective [8]. A text monitor checks the artifact the optimizer is free to shape.
On their own terms the documents do not conflict, because the recommendation Pachocki lands on already contains the brake [12]. What the material measures is the other side of it. August 2026 was the all-time high for experiments per active experimenter, across twenty months of tracking [6][16]. Whether the brake has ever been pulled is not a quantity in this data.
Ranked by verification strength, evidence, and original report placement.
OpenAI published two documents on September 6: one showing how quickly AI agents are taking over pieces of its own research process, and one by Chief Scientist Jakub Pachocki.
In his document Pachocki argued that this progress could lead to recursive self-improvement, while warning that alignment and monitoring may not be ready for what comes next.
The Neuron frames the two documents as a stark tension: OpenAI is getting better at using AI to build better AI, while its own chief scientist says no lab has solved the safety problem well enough to keep scaling at maximum speed for much longer.
OpenAI says it has now reached the 'automated research intern' milestone it promised for September 2026, meaning an AI system that can complete well-defined research tasks under human direction, including work that would take a skilled researcher a few days.
By mid-August, OpenAI researchers were using 3.1 agent-workdays of effort for every human workday.
The median researcher by agent usage was spending more than $600 per day of inference at API prices.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One issuer, one relay
Every figure in this story starts inside OpenAI and reaches the reader through a single summarized bullet list, with neither "agent-workday" nor "experiments per active experimenter" defined and no methodology attached. Pachocki's arguments fare better: they are quoted closely enough that a reader could check them against his essay. The throughput numbers have nothing to be checked against, and The Neuron's second posting is the first one again, so it adds no independent weight.
Deep inside one lab
The usage is real, dated and unusually granular for a frontier lab: a mid-August effort ratio, an August record for experiments per experimenter, a spend median, an intervention rate on multi-hour runs. It is also confined to OpenAI, counted by OpenAI, and drawn from researchers who already use agents, which is what the phrase "median researcher by agent usage" quietly concedes. Nothing here speaks to uptake anywhere else.
Self-graded milestone
The "automated research intern" milestone is scored against a target OpenAI set for itself, and the essay's harms are projections about superhuman intrusion, blackmail and engineered pathogens with no evaluation behind them. Pulling the other way, the checkable figures are modest about autonomy, and The Neuron keeps the human-intervention caveat in the same bullet list rather than burying it. That restraint is why the gap is small rather than wide.
Both documents serve the issuer
OpenAI gains from each half of this release: the throughput report advertises capability to customers and recruits, and the alignment essay casts its chief scientist as the person willing to say slow down. The cost detail carries the same fingerprint, since $600 a day is priced at the company's API rates, which is what OpenAI charges rather than what the inference costs it. Every party quoted in this reporting is OpenAI or its own summarizer, which leaves the numbers unchecked by anyone with an outside interest.
Thin corroboration
Dates, quotations and figures line up with one another, and Pachocki's essay material is specific enough to hold him to, so what OpenAI said is solid ground. What the numbers mean is shakier: one publisher, one source company, undefined metrics, and both of The Neuron's postings break off mid-sentence in the chain-of-thought section our coverage treats as the heart of the story.
product
OpenAI's median researcher spends more than $600 a day on inference at API prices1 publisher
leadership
OpenAI's chief scientist calls for mandated safety bars enforced from outside the lab1 publisher
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 publisher
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 publisher
Publishers with included, body-backed reporting in this cluster.
2 articles · September 6, 2026