Build1 publisher3 min readPublished
OpenAI's case for Codex as a general agent harness rests on one recycled tax pilot
Greg Brockman pitched Codex as infrastructure for non-coding products using 7,000 tax returns and a 31% time saving, both from a pilot OpenAI published months earlier with a partner it has a stake in.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Greg Brockman (@gdb), OpenAI's president and co-founder, pitched Codex on August 20th as infrastructure for products outside software development, pointing to a tax-preparation system that processed 7,000 returns and reduced accountants' preparation time by about a third; he argued developers can use Codex's open-source harness as the operating layer for specialized agents. Codex was originally put in terminals and development environments as a coding assistant.
- The figures describe a pilot from the 2025 tax year rather than a new deployment.
- OpenAI first detailed the project on May 27th, after six months of work with Thrive Holdings and Crete Professionals Alliance, the accounting network that rebranded as Current on June 2nd.
- The Codex repository is published under the Apache 2.0 license, and its App Server exposes the agent harness through a bidirectional interface that developers can embed in other products. The open-source code handles agent threads, tool execution, configuration and approvals; access to OpenAI's models still requires a ChatGPT account or API setup.
- Tax AI was built for Current's network of accounting firms. Participating accountants uploaded source documents and client notes, and the system extracted information and prepared submissions for tax-engine review. The pilot covered 1040 individual returns and 1041 returns for estates and trusts.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Greg Brockman, OpenAI's president and co-founder, used August 20th to pitch Codex as infrastructure for products outside software development, and the evidence he offered was a tax-preparation system that processed 7,000 returns and cut accountants' preparation time by about a third [1]. Those figures come from a pilot in the 2025 tax year that OpenAI had already detailed on May 27th, after six months of work with Thrive Holdings and the accounting network Crete Professionals Alliance, which rebranded as Current on June 2nd [2][3]; OpenAI took an ownership stake in Thrive Holdings in December 2025 [17].
The product claim is the interesting part. Codex began as a coding assistant inside terminals and development environments, and the repositioning is toward an operating layer for specialized agents [1][19]. The repository is Apache 2.0 licensed, and an App Server exposes the harness through a bidirectional interface developers can embed in their own products, handling agent threads, tool execution, configuration and approvals [5]. Access to OpenAI's models still requires a ChatGPT account or API setup [5]. That is a coherent business: give away the plumbing, meter the inference.
The tax numbers are thinner than they look. Tax AI was built for Current's firms, taking uploaded source documents and client notes, extracting information and preparing submissions for tax-engine review across 1040 individual returns and 1041 estate and trust returns [6]. OpenAI says data entry on medium- and high-complexity filings can run as much as eight hours per return [7]. Current reported an average 31% reduction in preparation time across the 7,000 returns [8], which on that eight-hour figure would be roughly two and a half hours, though the 31% is an average across the whole mix rather than the hard cases [1]. OpenAI also reported throughput up about 50% and drafts up to 97% accurate; Current later said as high as 98% [9]. Every one of those numbers comes from the two organizations that built and deployed the system, and the top-line accuracy figure does not say how results varied by return complexity [10].
The one measure with a shape to it is the improvement curve. At launch, a quarter of evaluated returns reached at least 75% correct field completion; six weeks later 86% did, while the system moved from W-2 and 1099 inputs into K-1 forms and rental-property schedules [11]. That is a 61-point gain [2], but the bar is 75% of fields, meaning a return can pass with a quarter of its fields wrong or missing [4], and 14% of evaluated returns still fell short of even that [3].
The mechanism is worth copying regardless of the marketing. Tax AI logged what it proposed, what the accountant changed and what entered the filed return; repeated errors were grouped into a finding, turned into a targeted evaluation, and handed to Codex as a bounded task with the production trace, source documents, expected tax-engine output, code and test commands [13][14]. Ambiguous cases went back to engineers rather than becoming code changes, because a correction can reflect judgment, a carried-over value or a change elsewhere in the filing process rather than a model error [14][15]. Codex's remit stopped at the extraction and mapping layer; humans kept architecture, product decisions, releases and final sign-off [12]. Rental-property support alone took about six weeks and substantial oversight to hit 90% precision and recall [16].
Watch for a vertical case study from a company OpenAI has no stake in, accuracy broken out by complexity band, and whether Current's next filing season reports volume rather than percentages.