Skip to content

Build1 publisher3 min readPublished

AppZen grades its own finance models ahead of frontier AI on five of six expense-audit controls

AppZen says its ZenLM Plus finance models led seven frontier models on five of six expense-audit controls in a test the company ran itself. Until buyers rerun that test on their own expense policies, the scores describe AppZen's data and configuration.

The Engineer · Build desk

Illustration accompanying AppZen grades its own finance models ahead of frontier AI on five of six expense-audit controls

What happened

  • ZenLM Plus lost on receipt verification, scoring 93.3 while Gemini 3.1 Pro and Sonnet 5 each scored 94.1.
  • AppZen's Expense Audit product uses the models to check customer rules, receipt validity, transaction details, itemization and duplicate claims across reports.
  • AppZen puts ZenLM Plus's modeled inference cost at about half of GPT-5.6 Luna's per 1,000 audited expense lines and about one-fiftieth of Opus 5's.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Routing is built into the product, so a buyer's evaluation has to score each control separately; an aggregate win cannot tell a team which model should own receipt verification.
  • cost The buyer's own pilot has to check the claim of one-fiftieth of Opus 5's cost, because AppZen has not laid out its full cost method.
  • capability A finance team can put ZenLM Plus on the controls it wins and keep a frontier model on the rest, so adopting it for part of the audit is a real option.

Each system in AppZen's comparison got the same expense data, supporting documents and customer configuration. Each was scored on precision, recall and the combined F1 [7]. The evaluation is AppZen's own [7]. The announcement does not give the test-set size, the prompts, the model settings or an independent evaluation [8]. Of those four, the prompts matter most. ZenLM Plus is a family of finance-specific models [1]. A frontier model sees a customer's categories and thresholds only through a prompt someone wrote for it, and in this test AppZen wrote it. So the 9.1-point lead over GPT-5.6 Sol on policy-category cases [1] compares AppZen's model with AppZen's prompting of the seven frontier models [2].

For that lead to transfer, a buyer's expense data has to look like the test set. The models draw on fields employees enter, receipt extractions, merchant and card data, and each customer's configured categories and thresholds [12]. One case AppZen cites is separating a payment slip that shows a card was charged from a merchant receipt that substantiates what was bought [13]. A company whose staff mostly photograph card-terminal slips has a different document mix from one that submits itemized hotel folios. The 9.1-point win and the 0.8-point loss on receipt verification [3] come from the same test set of unstated size. Neither can be read with a margin of error.

The receipt-verification loss is the most credible result in the release. Gemini 3.1 Pro and Sonnet 5 each beat ZenLM Plus there [6], and AppZen designed for that outcome. Its Mastermind Platform routes work among its specialized models, deterministic rules and frontier models, then applies the output inside finance workflows and controls [14]. A router is the right design for a model that wins most controls and loses some. According to RuntimeWire, Kale, the chief executive, described that blended approach as part of the product, and Verma, the CTO, said the platform can use a frontier model for a task where it is the better fit [15].

The cost claim uses different comparison models from the accuracy claim. The strongest frontier accuracy result came from GPT-5.6 Sol [4]. The two cost comparisons AppZen gives are against GPT-5.6 Luna and Opus 5 [9]. By AppZen's own numbers, Opus 5 scored 85.2 on non-conforming receipt detection [2], against 92.4 for ZenLM Plus [5]. A team that routes receipt verification to a frontier model would be paying for Gemini 3.1 Pro or Sonnet 5, the two models that won that control [6].

In my view, the pilot that settles this runs one control at a time, on expense lines the team has already audited. The frontier comparison in that pilot should use a prompt the buyer's own staff wrote. RuntimeWire makes a similar point: customer deployments will test whether the claimed accuracy and cost advantages hold on real expense policies [16].

What to watch

  • An independent or customer-run evaluation that publishes test-set size, prompts and model settings for the ZenLM Plus comparison.
  • Whether AppZen publishes its cost methodology, including figures for Gemini 3.1 Pro and Sonnet 5, the models that won receipt verification.
  • Which controls Mastermind sends to frontier models in early customer deployments.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories