Build2 publishers3 min readPublished
OpenAI reports 3.1 agent-workdays of agent runtime for every human workday, then spends much of the same report explaining why research did not get 3.1 times faster. The residue lands on review, compute allocation and deciding what to run.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Multiply the ratio out. OpenAI converts the time an agent spends on a task into standard eight-hour units, so 3.1 agent-workdays against an eight-hour human day is 24.8 hours of agent runtime per human day [4]. Researchers run several agents at once, which is why the figure describes how long agents were busy rather than what came back finished [3]. Runtime is the easiest thing in this stack to instrument, which is presumably why it is the number in the headline.
OpenAI makes the same point at greater length. Faster code and more experiments do not make the research process 3.1 times faster, because research also includes deciding what to pursue, designing experiments, analyzing results, allocating compute, catching failures and applying safety controls, and speeding up one stage can move the waiting line somewhere else [15]. The report adds that code output and experiment counts are easier to measure than their relationship to research progress, and that the least automatable tasks grow as a share of human work as automation improves [16].
The phase data agrees. Using Epoch AI's six-phase taxonomy of Decide, Design, Build, Run, Analyze and Communicate, OpenAI found activity rose in all six between January and August, while agents still did relatively little of the work of deciding what research to pursue [7]. Build got cheaper while Decide stayed where it was.
The intervention rate was produced by another model judging agent performance, over a period when success rates were improving [8]. For that number to transfer, the grader's notion of success has to match what your reviewers would accept, and your task mix has to sit in the same four-to-eight-hour band. OpenAI's own definition of the automated research intern is a well-defined task that would take a skilled person several days, with a human still in charge [9], which is a harder band than the one where humans were still stepping in.
Then the allocation side. In the week after the Astra access change, Astra-class GPU allocation fell 59.2%, allocation to other models rose 17.2%, and that rise covered roughly 85% of the drop [13]. If both percentages are measured against each pool's own prior week, the non-Astra pool was about 2.9 times the size of the Astra pool: 0.592 x 0.85 / 0.172 = 2.93 [14]. The work moved because there was somewhere roughly three times as large to move it into. Compute also rose significantly as experiment counts grew [18], and Astra's persistent agents let researchers hand off multi-day assignments, which adds to what one person has to track [19].
The published metrics are runtime, code volume, experiment counts, compute and spend; a completed-task count is not among them [22]. That leaves the relocation of the constraint well evidenced by OpenAI's own qualification, its July and August incidents and the intervention rate, and the size of the net gain unmeasured.
Ranked by verification strength, evidence, and original report placement.
On September 6, 2026, OpenAI published a detailed look at how coding agents are changing work inside its research organization.
By mid-August, OpenAI's agents were logging 3.1 agent-workdays for every human workday across the research organization, with coding-agent use climbing through 2026.
Using a taxonomy from Epoch AI, OpenAI split agent work into six areas (Decide, Design, Build, Run, Analyze, Communicate) and found activity increased across all six between January and August, although agents still did relatively little of the work of deciding what research to pursue.
OpenAI's report qualifies that faster code and more experiments do not automatically make the whole research process 3.1 times faster, because research includes deciding what to pursue, designing experiments, running them, analyzing results, communicating findings, allocating compute, catching failures and applying safety controls, and speeding up one stage can move the waiting line somewhere else.
OpenAI said code output and experiment counts are easier to measure than their relationship to research progress, and that as automation improves the least automatable tasks become a larger share of human work and may become the next constraint.
Agents wrote research and infrastructure code, monitored experiments and provided enough technical support that attendance at OpenAI's debugging office hours fell, prompting one team to stop holding the sessions altogether.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One issuer, self-graded
The 3.1 ratio, the $600 median, the 59.2% GPU drop and the more-than-half intervention rate all come from OpenAI describing its own research organization, and the success rates behind that intervention figure were scored by another OpenAI model rather than an outside evaluator. The two write-ups agree because they read the same document, so the agreement adds no verification. What keeps the score from falling further is how specific and dated the disclosure is, and that OpenAI itself publishes the caveat undercutting its headline number.
Deep, but inside one lab
Usage is real and unusually well quantified: runtime per human workday, per-researcher daily spend, GPU allocation moving between model pools within a week, and a debugging office hour that stopped being worth holding. All of it happens inside OpenAI, among researchers with unmetered access to their employer's frontier models, which makes it a poor proxy for anyone else's adoption. The one external trace is indirect — safeguards that developers may be meeting as unexplained API interruptions.
Headline oversells, issuer says so
A figure like 3.1 agent-workdays per human workday invites reading as a threefold speedup, and dev.to opens by noting it sounds like an extra Monday through Wednesday installed inside Monday. But that figure is runtime converted into eight-hour units, from agents run in parallel, and no completed-task count appears anywhere in the record. The gap stays modest because OpenAI supplies the correction in the same report and both publishers led with it -- The New Stack's headline argues the agents create more work, not less.
Vendor measuring its own product
OpenAI is reporting on the usefulness of the agents it sells, quoting internal consumption at its own API list prices, and dating its progress against a self-defined milestone with a self-defined successor due March 2028. The dev.to piece has a smaller interest of its own: it uses the report as a lead-in to the author's free prompt pack. Against that, the disclosure includes material a promotional document would omit — agent-caused outages, a safety threshold that forced a model behind higher-security walls, and an explicit statement that the countable metrics do not show progress.
Consistent, unverifiable
The figures are precise, dated and consistent across both accounts, and the caveats come from the party with the least reason to volunteer them, which is why this sits above the middle. It cannot go higher while every number has a single origin, Astra is a model no outsider can test, and the grading of agent success was done in-house.
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 publisher
science
OpenAI declares Astra the first model to reach its Critical cyber threshold1 publisher
leadership
Every notable AI release today arrived with a grade written by its own vendor1 publisher
build
OpenAI slows training after its own model breached Hugging Face: a safety gate builders must plan for7 publishers
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026
1 article · September 7, 2026