Build1 distinct publisher3 min readUpdated
Greg Brockman pitched Codex as infrastructure for non-coding products using 7,000 tax returns and a 31% time saving, both from a pilot OpenAI published months earlier with a partner it has a stake in.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
Greg Brockman, OpenAI's president and co-founder, used August 20th to pitch Codex as infrastructure for products outside software development, and the evidence he offered was a tax-preparation system that processed 7,000 returns and cut accountants' preparation time by about a third [1]. Those figures come from a pilot in the 2025 tax year that OpenAI had already detailed on May 27th, after six months of work with Thrive Holdings and the accounting network Crete Professionals Alliance, which rebranded as Current on June 2nd [2][3]; OpenAI took an ownership stake in Thrive Holdings in December 2025 [17].
The product claim is the interesting part. Codex began as a coding assistant inside terminals and development environments, and the repositioning is toward an operating layer for specialized agents [1][19]. The repository is Apache 2.0 licensed, and an App Server exposes the harness through a bidirectional interface developers can embed in their own products, handling agent threads, tool execution, configuration and approvals [5]. Access to OpenAI's models still requires a ChatGPT account or API setup [5]. That is a coherent business: give away the plumbing, meter the inference.
The tax numbers are thinner than they look. Tax AI was built for Current's firms, taking uploaded source documents and client notes, extracting information and preparing submissions for tax-engine review across 1040 individual returns and 1041 estate and trust returns [6]. OpenAI says data entry on medium- and high-complexity filings can run as much as eight hours per return [7]. Current reported an average 31% reduction in preparation time across the 7,000 returns [8], which on that eight-hour figure would be roughly two and a half hours, though the 31% is an average across the whole mix rather than the hard cases [1]. OpenAI also reported throughput up about 50% and drafts up to 97% accurate; Current later said as high as 98% [9]. Every one of those numbers comes from the two organizations that built and deployed the system, and the top-line accuracy figure does not say how results varied by return complexity [10].
The one measure with a shape to it is the improvement curve. At launch, a quarter of evaluated returns reached at least 75% correct field completion; six weeks later 86% did, while the system moved from W-2 and 1099 inputs into K-1 forms and rental-property schedules [11]. That is a 61-point gain [2], but the bar is 75% of fields, meaning a return can pass with a quarter of its fields wrong or missing [4], and 14% of evaluated returns still fell short of even that [3].
The mechanism is worth copying regardless of the marketing. Tax AI logged what it proposed, what the accountant changed and what entered the filed return; repeated errors were grouped into a finding, turned into a targeted evaluation, and handed to Codex as a bounded task with the production trace, source documents, expected tax-engine output, code and test commands [13][14]. Ambiguous cases went back to engineers rather than becoming code changes, because a correction can reflect judgment, a carried-over value or a change elsewhere in the filing process rather than a model error [14][15]. Codex's remit stopped at the extraction and mapping layer; humans kept architecture, product decisions, releases and final sign-off [12]. Rental-property support alone took about six weeks and substantial oversight to hit 90% precision and recall [16].
Watch for a vertical case study from a company OpenAI has no stake in, accuracy broken out by complexity band, and whether Current's next filing season reports volume rather than percentages.
Ranked by verification strength, evidence, and original report placement.
Greg Brockman (@gdb), OpenAI's president and co-founder, pitched Codex on August 20th as infrastructure for products outside software development, pointing to a tax-preparation system that processed 7,000 returns and reduced accountants' preparation time by about a third; he argued developers can use Codex's open-source harness as the operating layer for specialized agents. Codex was originally put in terminals and development environments as a coding assistant.
The figures describe a pilot from the 2025 tax year rather than a new deployment.
OpenAI first detailed the project on May 27th, after six months of work with Thrive Holdings and Crete Professionals Alliance, the accounting network that rebranded as Current on June 2nd.
The Codex repository is published under the Apache 2.0 license, and its App Server exposes the agent harness through a bidirectional interface that developers can embed in other products. The open-source code handles agent threads, tool execution, configuration and approvals; access to OpenAI's models still requires a ChatGPT account or API setup.
Tax AI was built for Current's network of accounting firms. Participating accountants uploaded source documents and client notes, and the system extracted information and prepared submissions for tax-engine review. The pilot covered 1040 individual returns and 1041 returns for estates and trusts.
OpenAI said data entry for medium- and high-complexity filings can consume as much as eight hours per return, including pulling information from prior-year filings, spreadsheets and other inconsistent client documents and mapping it to the correct tax fields.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-source and self-reported
The mechanism account is unusually specific — licensing terms, harness scope, the correction-to-evaluation pipeline, per-schedule precision and recall — but it rests on one publisher relaying figures produced by the two organizations that built and deployed the system. The source itself notes the metrics are self-reported and not broken down by complexity, and the two parties quote different accuracy ceilings.
One real production pilot, one vertical
This is a genuine production deployment at non-trivial scale — 7,000 1040 and 1041 returns through a network Current says now spans 48 firms and 2,000-plus employees — but it is a single vertical inside one partner network, and the harness-as-general-infrastructure claim has no disclosed third-party adopters. Expansion into bookkeeping, audit and IT help desk is stated as plan, not deployment.
General claim, single-vertical proof
The pitch generalizes from one recycled accounting pilot to Codex as the operating layer for vertical agents generally, while the same reporting shows the hard parts were domain-specific: weeks of oversight per schedule, human sign-off on every return, ambiguous cases escalated to engineers, and a self-reported accuracy ceiling that the two parties state differently. Positive but not extreme, because the source discloses the caveats and the underlying deployment is real.
Vendor holds equity in the deployment chain
Every number comes from OpenAI or Current, and OpenAI took an ownership stake in Thrive Holdings — with embedded staff and accounting named as a first target — before the results were promoted. The messenger is OpenAI's president pitching a company product, and the partner network has its own commercial interest in an AI-augmented service story.
Mechanism clear, magnitudes unverified
Confidence is moderate: the qualitative architecture — open harness, bounded Codex tasks, correction-driven evaluations, human approval — is described consistently and with enough specificity to be actionable, and the source flags its own limits. Confidence is capped by the single publisher, the self-reported and internally inconsistent metrics, and the absence of cost, compliance and complexity-level detail.
security
The AI security line item to fund first is log coverage, not another agent2 distinct publishers
security
OpenAI's Computer History writes a plaintext log of the workday. Decide before staff opt in.1 distinct publisher
product
The AI-wrote-it claim died in eight hours. The Actions injection pattern did not.1 distinct publisher
build
Waku 0.1.0 bets the product is the control plane, not another coding agent1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026