Build1 distinct publisher3 min readPublished
The two-week Enzyme removal is credible enough, but it is priced against a five-year staffing plan written by people who did not want the project, and Airbnb's published numbers let you check the rate that plan assumed.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Enzyme and React Testing Library are not two dialects of the same test. An Enzyme test operates on a component instance; RTL asserts against the rendered DOM, so the test sees the whole rendered page rather than the component it was written for [6]. Assertions have to be re-expressed in terms of what a user can see, and for complex user journeys the two approaches diverge sharply [7]. Asana went from the first to the second [4]. Both sides of the comparison are describing genuinely awkward work.
The units are where it comes apart. The arithmetic behind the $6M is four engineers at circa $300K a year for five years, according to Gergely Orosz's reading of the case study [5], and it multiplies out cleanly to the published figure [15]. That is also about 1,040 engineer-weeks of budgeted labour [16]. The other side of the comparison is 1.5 weeks of engineering effort [3], and the write-up does not say how many engineers those 1.5 weeks cover, which is a curious omission in a comparison this lopsided.
The $12K is model and infrastructure spend, not the cost of the work [1]. Charge the engineer time at the salary the $6M itself assumes and you add roughly $8,650, for about $20,700 all in [17]. Against an advertised gap of $5.99M [20], that is noise either way.
Airbnb's numbers let you test the rate the plan assumed. Its manual estimate was 1.5 engineering years for 3,500 Enzyme component test files [8]. At 250 working days a year, that implies about nine files per engineer-day [18], the top of the five-to-ten range Orosz guesses Asana's planners used [13]. Push Asana's implied 20 engineering years through Airbnb's rate and the plan is sized for roughly 46,700 test files [19]. If Asana's suite is nowhere near that, the five years measured appetite rather than throughput. Orosz puts it plainly: you estimate a project like this when it is work you really do not want to do [14].
Airbnb also shows where the effort actually sits. Its team built loops that retried failed migrations, and once those existed 75% of files went through in four hours [9]. The remaining quarter needed a purpose-built refactor pipeline, which took the total to 97% over four days, with engineers finishing the last 3% by hand over a week [10]. That was March 2025, on Claude 3.7 Sonnet [11]. The bulk was hours; the tail was a week of humans.
Orosz still finds two weeks credible for Asana, a year on from Airbnb [12]. The outcome can stand while the denominator gets audited. Whether a saving like this transfers depends on who authored the counterfactual and whether they wanted the work, the unit both figures are denominated in, and what the smaller figure leaves out. Calendar weeks set against engineer-years fails on the second before you get to the third.
Ranked by verification strength, evidence, and original report placement.
OpenAI's case study states the previous staffing plan for the work was expected to take at least five years and was estimated to cost roughly $6 million.
OpenAI's case study states that Asana used OpenAI Codex to remove Enzyme, an outdated testing system, in two weeks, with model and infrastructure costs totalling about $12,000.
Per OpenAI, Enzyme was fully removed after 1.5 weeks of engineering effort spread across two calendar weeks.
Asana migrated from Enzyme to React Testing Library, which the Pragmatic Engineer newsletter says would not be a simple migration.
According to Gergely Orosz, OpenAI's arithmetic rests on an estimate of four engineers, each on circa $300K a year, spending five years on the project.
Enzyme is oriented towards component testing and operates on a component instance, while React Testing Library operates on the rendered Document Object Model, so the test sees the whole rendered page rather than only the component it was written for.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Codex learns to click: the coding agent stops typing patches and starts operating the machine1 distinct publisher
build
A build-tool swap that took 70 days and 166 files, and the build was the easy part1 distinct publisher
product
OpenAI's plan to hand everyone a coding agent leaves the hard part to the model1 distinct publisher
build
Twenty-three security checks, zero coverage: AI coding agents as build-pipeline attack surface1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Realized figures documented, counterfactual unverified
The realized side is specific and self-consistent: a quoted vendor case study with a dollar figure and an effort figure, plus Airbnb's independently published phase-by-phase migration numbers that make the timeline plausible. The comparison side is weak: the five-year, $6M baseline is a self-reported internal estimate, no Asana test file count is disclosed, and the entire cluster rests on one publisher with no independent corroboration or vendor response to the inflation critique.
Two named production migrations, one vendor-disclosed
Adoption is real and named rather than aspirational: two large engineering organizations completed Enzyme test-suite migrations with LLM assistance, roughly a year apart, with concrete file counts, phase timings and costs. It remains two data points from one narrow workload class (test framework migration), one of them disclosed by the model vendor itself, with no evidence in the supplied material of broader diffusion.
Savings headline overstated; capability claim roughly fair
The capability claim - that a migration once treated as impractical was done in weeks - is well supported and, if anything, understated in the cluster. The financial claim is inflated. The $5.99M headline is the difference between a self-reported five-year staffing plan and a $12,000 bill that omits the 1.5 engineer-weeks actually spent; adding that time at the case study's own $300K assumption yields about $20,700. The plan itself equals about 1,040 engineer-weeks and, at Airbnb's published rate, would imply roughly 46,700 test files, a scale the source never evidences for Asana.
Vendor authors both the savings math and the baseline framing
Incentives run strongly in one direction on the headline figure: OpenAI publishes the case study, quantifies the savings attributable to its own product, and sources the counterfactual from a customer estimate. The article adds a second incentive layer on the baseline itself - that five-year plans get written for projects engineers do not want to do, which biases the estimate upward before any vendor touches it. The commenting publisher runs a paid subscription newsletter, a mild incentive toward contrarian framing of vendor claims, and it discloses that it contacted Asana directly.
Single publisher, verifiable arithmetic, unverifiable baseline
Confidence is limited by cluster structure more than by internal quality. One publisher supplies every fact; the quoted case study text and Airbnb's figures are specific and the derived arithmetic reproduces cleanly, which supports the realized-cost and capability conclusions. But no independent source confirms the $6M plan, Asana's suite size is unknown, the article's Asana clarification section is truncated in the supplied body, and neither OpenAI nor Asana answers the inflation charge.