Skip to content

Build1 publisher2 min readPublished

Seven finance experts wrote the 537 SEC-filing tasks that hold o3 to 46.8 percent

The Finance Agent Benchmark scores agents on recent SEC filings using Google Search and EDGAR access, and its best reported result is OpenAI's o3 at 46.8 percent accuracy and $3.79 a query. A full pass runs about $2,035.

The Engineer · Build desk

Illustration accompanying Seven finance experts wrote the 537 SEC-filing tasks that hold o3 to 46.8 percent

What happened

  • The Finance Agent Benchmark ships 537 expert-authored questions organised under a taxonomy of nine financial task categories.
  • The evaluation runs through an agentic harness that gives the model tools including Google Search and EDGAR database access.
  • OpenAI's o3 was the best-performing model at 46.8 percent accuracy, at an average cost of $3.79 per query.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Each correct answer costs about $8.10 at the paper's figures, and the review time for the roughly half that come back wrong is paid by whoever signs the output.
  • constraint A ceiling below 50 percent keeps a human in the loop on every answer, and the authors say further advances are needed before reliable deployment in high-stakes finance settings.
  • decision Teams quoting a score now have to declare their tool surface, because a licensed data feed or a pre-indexed filing store makes the same questions a different test.
  • capability Expert reasoning trajectories let a team see which step of an answer went wrong instead of only whether the final number matched.

The average cost per query is $3.79 and the set holds 537 questions, so one full pass costs $2,035.23 [5][1][11]. At 46.8 percent accuracy that pass returns about 251 correct answers and about 286 wrong ones [12]. You will run it more than once.

Whether 46.8 percent transfers depends on the tool surface. In this harness the agent gets Google Search and EDGAR database access, and the authors describe that toolset as sufficient to produce an accurate response [4]. An agent that also holds a licensed data feed or a pre-indexed filing store is taking an easier test, and its score is not comparable. Give it less than search plus filings and it will do worse. The $3.79 is an average across queries, so the questions that need many filing fetches cost more than the ones answered in a single lookup. If you report a number on this benchmark, report the tool list beside it.

The construction is the part I would trust. Seven domain experts from banks, hedge funds and private equity firms defined the taxonomy of nine task categories [2]. Experts wrote the questions, the answers and the step-by-step reasoning trajectories, every question is verifiable against a public filing in the SEC's EDGAR database, and each one went through peer review by the experts [3]. Table 1 holds the taxonomy, and the nine categories are not named in the abstract or the introduction [14].

The worked example of the failure it wants caught is specific: a model misinterprets an earnings report, incorrectly identifies a consistent earnings beat over four consecutive quarters, and an investment decision follows the error [10]. That is catchable only because each answer key is tied to a filing someone can open. On the state of the field, the authors wrote that the result "underscores the need for further advancements before reliable deployment in high-stakes finance settings" [6].

The motivation cited is time. Entry-level finance professionals sometimes spend up to 40 percent of their workweek gathering data instead of analyzing it, according to a figure the paper cites [9]. The questions span plain information retrieval up to complex financial modeling [7], so a score below half is spread across both the fetching and the modeling. Earlier benchmarks, the authors argue, miss the interactive environments and nuanced reasoning that industry-specific tasks demand [8].

What to watch

  • Per-category results from Table 1 would show whether failures cluster in retrieval or in the modeling tasks.
  • A rerun with cheaper or newer models would move both the $3.79 per-query average and the cost per correct answer.
  • Whether other teams post scores on this same harness or substitute their own data feeds, which would break comparability.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories