Build1 publisher3 min readPublished
Frontier data agents average 59.5 on Argo-Bench's 235-table warehouse tasks
Frontier models average 59.5 on Argo-Bench and clear 95 on only 34.8% of its 210 enterprise data tasks. The benchmark grades the bans and refunds an agent files against a hidden simulator, a step outside what text-to-SQL scores measure.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The benchmark simulates a New York City food delivery platform with 81 million orders in 2024, including pricing, fraud such as promo abuse, and courier and restaurant incentives.
- That simulator exports to a 235-table warehouse modeled on Oracle E-Business Suite, holding 7.5 billion rows.
- The simulator's ground-truth state is withheld from the warehouse, so agents must reconstruct facts from incomplete or denormalized data before they act.
- Agents finish each task by filing an action such as a ban list, budget allocation or refund batch, and the grader scores its outcome in the simulator instead of comparing SQL.
- According to the write-up, the strongest models fail most often at state reconstruction, treating the warehouse as authoritative when it is only a partial view.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A Spider or BIRD score covers the query-writing step alone, so it cannot vouch for an agent on the reconstruction and action steps where Argo-Bench's strongest models fail.
- exposure Once an agent holds write actions such as account bans or refunds, trusting a partial warehouse produces wrong bans and wrong refunds that land on real customers.
- cost Teams pay for harness work beyond the model: memory of explored tables, storage for intermediate results, hypothesis tracking, and staging each action for validation before submission.
- capability Per-query execution logs, step-by-step state snapshots and simulator diffs against the reference action let a team trace a wrong ban list back to the query where reasoning diverged.
Start with the write-up's example task: "Ban accounts that placed more than 10 orders from different zip codes in 24 hours" [14]. On a single-table benchmark that is one GROUP BY with a HAVING clause. In Argo-Bench the agent has to find the order history, payment method, device fingerprint and geolocation tables, work out how they join, and infer the fraud pattern on its own [15]. There is no fraud_flag column to select [15]. With one, the benchmark would take an afternoon to pass.
The write-up says a typical task means exploring 10 to 15 tables and running 3 to 5 exploratory queries. Then come aggregations such as percentiles, moving averages and cohort analysis, and only after that is an action drafted [13]. Each step feeds the next. A wrong assumption about one join carries into the aggregate and then into the ban list.
The schema is spread thin on purpose. Dividing the warehouse's 7.5 billion rows by the simulator's 81 million orders gives about 93 rows per order [1]. That ratio counts every table, so it overstates how far one order fans out. It still fits the layout the write-up attributes to real warehouses, where one transaction is spread across dozens of normalized tables [11]. Public datasets, by the same account, fit each business event into one table [11].
I would copy the grading design. Spider and BIRD give the agent a question, a schema and a database, then compare its output to a reference answer [9]. According to the write-up, audits of popular benchmarks find incorrect ground truth, so an agent marked as failing may in fact be right [10]. Argo-Bench ships an executable reference solution for every task, and each one uses only the warehouse the agent sees [7]. That proves each task can be solved from the agent's side and gives the grader a correctness baseline [7].
If the share is taken over all 210 tasks, 34.8% is about 73 tasks scored above 95 [2]. The other 137 or so fall short of that bar [3]. The results are reported for frontier models as a group [1]. The write-up does not set them beside the same models' Spider or BIRD scores. So the drop from query generation to stateful warehouse work follows from the design, but it is not measured as a gap [18].
For the 59.5 average to describe a team's own agent, that team's setting has to match the benchmark's [1]. The warehouse has to be a normalized ERP export in the Oracle E-Business Suite mould [2]. The work has to end in a filed action that is graded by its outcome [5]. And the agent has to be working from a partial view of the truth, as it is here [4].
What to watch
- Per-model scores in the full Argo-Bench results, to show whether the 59.5 average hides one model well above the group.
- A same-model comparison of Spider or BIRD scores against Argo-Bench scores, the measurement that would size the drop from query writing to warehouse work directly.
- Whether the simulator and 235-table warehouse become available for teams to run their own agent harnesses against.