Build1 publisherNot yet confirmed elsewhere3 min readPublished
OpenAI and Ironclad use production contract workflows to train and test computer-use agents
OpenAI and Ironclad published a case study on turning production contract workflows into training data and evals for computer-use agents. Scores from such an eval carry over to other software only as far as it shares contract approval's roles and branching.
The Engineer · Build desk

What happened
- Ironclad's flows are stateful: a contract can move from draft to legal review, back to sales, then to finance, with different permissions and screens at each step.
- OpenAI gets Ironclad's workflow graphs, multi-role approval chains, document versioning and redlining, external integration points and real production failure cases.
- Customer contracts cannot be used, so Ironclad has to generate synthetic contracts from templates and draw redline patterns from anonymized edits.
- Reproducible runs need snapshot environments with fixed UI state, versioned workflow definitions and deterministic approvals with no human in the loop.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A pass on this eval is a pass against fixed UI state and scripted approvers, so it predicts little about an agent facing a live reviewer or a redesign shipped last week.
- decision Teams buying computer-use agents have to judge whether their own software shares contract approval's shape before treating an Ironclad-based score as evidence about it.
- cost Building an equivalent eval in-house means logging every step and paying people to classify failures, since the post says that labeling is not automated.
- exposure Ironclad customers' anonymized edits become the source of redline patterns in a model lab's agent training and eval data.
Most of the public detail on this partnership comes from a dev.to breakdown, and its author hedges. The pipeline is presented as "the likely flow" and the snapshot code as a "plausible schema" [13][1]. The post does not include pass rates, task counts or the models tested. It is one engineer's reconstruction of what an eval built on the case study would need [2].
Ironclad's flows are worth testing on for two reasons. They hold state, and different users take different paths through them [4][6]. Fixing the UI state and scripting the approvals [11] removes much of both. Scripted approvers are easier to grade than lawyers. What remains is structure: the workflow graph, the role boundaries, the screens each role sees, and failure cases taken from production [5]. That is still more than a simulated browser task, which the post says is what most computer-use benchmarks are [8].
The sketched grader has a real flaw. Call validate_agent_run and it walks the expected steps by index. It returns False if the agent's trace runs short. It returns False at the first step that does not match. It checks the success criteria only after every step has matched [1]. An agent that completes the contract by a different valid route therefore fails [7]. The same post says human users take different routes [6]. A grader that scores the end state against the success criteria, and treats the step list as one reference path, avoids this. Because the schema is the author's guess, Ironclad's real grader may already do so.
Version handling is the strongest part. Ironclad ships product updates, and a UI change breaks the eval. The post says the infrastructure has to track product version alongside eval version, keep older snapshots backward-compatible, and flag test cases a UI change invalidates [12]. I think this is the right design. When a pass rate drops, a version-pinned eval can tell a worse agent apart from a moved button.
The post's broadest claim is about transfer: "If an agent can navigate this, it can navigate most SaaS tools" [9]. For that to hold, the target tool needs contract approval's shape, with role-scoped permissions, conditional branches and approvals that move between teams [3][4]. Software without that shape learns little from an Ironclad score. The only workflow that settles the question for a given team is its own. This pipeline shows the cost of building that test. The logger has to capture DOM snapshots, API request and response pairs, intent signals and success criteria at every step [14]. People then classify each failure, a step the post says is not automated [15]. The evidence for a wider move away from synthetic benchmarks is this one partnership, described secondhand [2][13].
What to watch
- Whether OpenAI or Ironclad publish pass rates, task counts or the snapshot format itself, which would let outside teams test the post's transfer claim.
- How the eval handles Ironclad's next UI release: whether invalidated cases are flagged and replaced, or the test set shrinks.
- Whether other SaaS vendors open their workflow engines to model labs on similar terms.