Leadership1 distinct publisher3 min readPublished
Up to four coding agents finished in roughly a week and a half a test-suite migration Asana had been grinding through since 2022, which makes the queue of someday cleanups less a matter of capacity than of what a leader schedules first.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The compression Asana describes is conditional, and the conditions are the part that transfers to anyone else's backlog. By the company's own account, the codebase already held most of the context: the earlier decision to adopt React Testing Library, well-designed test helpers, clear conventions, and real examples an agent could read without being told they existed [9]. The job also had a definition of done a machine could check, no Enzyme anywhere, plus fast feedback from typechecking, linting, tests and CI [10]. The years of staffed, incremental progress that preceded the sprint [3] are what made the agent run legible, which is a less quotable finding than the two-week number.
The ratio is worth doing by hand. Five remaining years is about 260 weeks, so a week and a half of engineering time divides into that projection roughly 173 times [1]. A skeptic will point out that the units differ, and the skeptic is correct: the five years was calendar time at whatever staffing product teams could spare around roadmap work, while the week and a half was concentrated engineering time with up to four agents running through the night [5]. Discount the ratio by an order of magnitude and the management question still changes, because what kept the migration in the queue was never the total hours. It was the unwillingness to book engineer-years against cleanup in any single quarter [3].
The constraint that justified deferral was capacity, and what replaces it is order. If a bounded migration with a verifiable finish line costs a sprint rather than several multi-engineer-year projects [3], the useful question for an engineering leader this quarter is which quietly abandoned items have that shape, and which would have to be reshaped before they do. Asana names its own list as long-running migrations, rewrites and performance problems it had assumed would always take years [15]. Note also what did not help: tracked tickets, running notes files, spawned sub-agents and a detailed conventions prompt all made the output worse than five sentences did [7], which is consistent with the claim that the environment, not the instruction set, was carrying the knowledge.
The bill does not disappear; it moves. Asana reports the agents were rarely the bottleneck and its own tooling was, citing a lint step that sometimes ran past ten minutes and mismatches between CI and local checks [12]. Internal docs that still recommended Enzyme steered the agent wrong, which the company calls misleading training material for every agent that reads it [11]. Documentation hygiene and CI latency are the classic unfunded complaint, because the cost has always fallen on humans who could route around it. Agents route around nothing, and they meet the same ten-minute lint step at machine frequency.
The honest limit of this record is scope: one migration, at one company, on a task with a binary done condition. Asana's post begins to qualify that not every deferred problem will collapse the same way, and the text available to us stops mid-sentence [16]. Unresolved too is what this does to the people whose work it absorbs, since the post acknowledges that engineering identity is bound up in writing code by hand and that the change can feel unsettling [15]. A manager who hands a five-year migration to four overnight agents is also editing what the team believes it is paid for, and that conversation arrives in the same quarter as the schedule.
Ranked by verification strength, evidence, and original report placement.
The published text supplied ends mid-sentence as Asana begins to qualify the result, reading: "Not every one of those problems will collapse from years to".
In 2022 Asana set out to migrate its frontend test suite off Enzyme, an aging testing library, and onto React Testing Library (RTL).
Asana says Enzyme had lost community support, did not work well with newer versions of React, and encouraged tests tightly coupled to implementation details rather than user-visible behaviour.
Asana set a goal of finishing the whole migration in one week; it took about a week and a half of engineering time, and Enzyme is now gone from the codebase entirely.
Asana used OpenAI's Codex with frontier models on extra-high reasoning, running up to four agents at a time, each pointed at a different directory, keeping the machine from sleeping so agents ran through the day and overnight, with check-ins each morning and evening to review progress and open pull requests.
The entire prompt was five sentences: migrate the repo from Enzyme to RTL-style tests, follow existing norms and best practices in the codebase, migrate all files in a given directory, test changes with a named test command, and bias for migrating easy-to-convert files first.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
OpenAI's Asana case study prices 1.5 weeks of engineer time against a five-year staffing plan1 distinct publisher
build
Codex learns to click: the coding agent stops typing patches and starts operating the machine1 distinct publisher
product
Four leaderboards, four denominators: what you buy when you standardize on a coding agent1 distinct publisher
build
Codex can now ask and keep going, which deletes the only checkpoint you were getting for free1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-party, unusually specific, entirely unchecked
Asana shows more of its work than most companies do — the prompt verbatim, the spend to the thousand, the setups that failed, the lint step that ran ten minutes — and that specificity is worth something. But every load number in the story comes from the same author with no outside look: the five-year runway is an internal extrapolation, the $6M is called back-of-napkin on the page, and the volume that actually moved is never stated in files, tests or pull requests. 'Enzyme is now gone' is checkable in principle and by nobody so far.
Shipped in production, sample size one
This is not a pilot or a demo: a real frontend suite at a real company, with a binary outcome someone inside can falsify, plus a dollar figure attached. That puts it well above the usual agent anecdote. What it does not have is repetition — one codebase, one task shape, one team, and no second migration yet to show the harness investment travels the way Asana says it will.
The multiple is doing more work than the measurement
Years-to-a-sprint framing compares apples with hours: a partly staffed five-year projection against a week and a half of hands-on time, and the same page's two-calendar-week figure already pulls the multiple down by a third. The $12K-versus-$6M line stretches further, since the $6M scope quietly includes coverage work and legacy cleanup the migration was not originally chartered to do. Asana pulls itself back — not every deferred problem collapses, and the friction it names is mostly its own — which is why the gap is a lean, not a chasm.
A customer story written inside the vendor's partnership
The last line of the post states plainly that it is part of an ongoing Asana–OpenAI collaboration exploring what Codex can take on, so the two parties best served by a large multiple are the ones producing and hosting the number. Asana also has its own reason to be seen moving fast on agents. The disclosure is honest and it is at the end, after the reader has absorbed the ratio; the failure modes admitted along the way are real, but they are all failures the company has already fixed.
Trust the mechanics, hold the multiple loosely
We are fairly sure of the how — the prompt, the parallel agents, the stale docs, the slow lint — because those details are concrete, self-incriminating, and consistent within one account. We are much less sure of the how much, because the scope moved is never quantified, the baseline and the manual-cost estimate are both the author's own arithmetic, and no independent party has looked. Add the partnership disclosure and this sits just above the midpoint.