Skip to content

Invest1 publisher3 min readPublished

Rippling graded 2,100 agent runs per model. The cheap one basically tied the flagship.

A production test across 15 models put seven of them inside a one-point spread on pass rate. On constrained payroll work, the price premium bought speed, not correctness.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Rippling's President and CPO Matt MacInnis published a test of AI models doing real work inside a real production system, rather than a leaderboard or invented tasks.
  • The study covered about 2,100 graded attempts per model, across 15 models, on actual personnel, payroll, and financial records.
  • Seven models landed between 88.5% and 89.5% pass rate, a one-point spread; adding the leader at 91.0% makes the whole competitive group 2.5 points wide.
  • MacInnis writes that defaulting to whatever the big lab just shipped means paying a 3x to 7x premium for a difference customers cannot detect.
  • Every attempt either passed Rippling's production correctness checks or failed, with no partial credit; an attempt that never finished counted as a failure.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Rippling's president and chief product officer, Matt MacInnis, published what most B2B companies keep internal: roughly 2,100 graded attempts per model, across 15 models, run against actual personnel, payroll and financial records [1][2]. The consequence for anyone signing an inference bill is that seven models finished between 88.5% and 89.5% pass rate, a one-point spread, while the prices attached to that band differ by up to seven times [3][4].

The grading is why the result carries weight. Every attempt either passed Rippling's production correctness checks or failed, with no partial credit, and an attempt that never finished counted as a failure [5]. According to MacInnis, that is a harder grader than most published benchmarks use [6]. The work was narrow and operational: headcount by department, tenure distribution, raising base salaries for everyone meeting a condition by 10%, onboarding a hire through a checklist, scheduling a termination and routing its approvals, entering amounts into a pay run from a spreadsheet [7].

On those tasks, 15 models collapse to three defensible purchases. Opus 4.6 leads accuracy at 91.0% for $1,453 per full test run, with 154 seconds at the slowest 10% [8]. GPT-5.5 at medium effort scored 89.5% for $1,435, an $18 gap [9] that is about 1.2% of the bill [1]. GPT-5.5 at low effort scored 88.8% for $1,308 with a 130-second tail [10]. GLM 5.2 scored 88.7% for $621 with a 243-second tail [11]: 57% cheaper than the accuracy leader for 2.3 points of pass rate [2], and 53% cheaper than GPT-5.5 low for a tenth of a point, at the cost of 89 extra seconds at the tail versus Opus 4.6 [3][4].

The expensive end does not survive that comparison. Fable 5 was the most expensive model tested and finished fifth, matching GPT-5.5 low on accuracy at 3.3 times the price and twice the wait [12][13]. Opus 5 was beaten on both price and speed by Grok 4.5, which scored 0.2 points behind it for 68% less [14]. Opus 4.6 beat both newer Anthropic models on accuracy while costing 42% less than Opus 5 and a third of Fable 5 [15]. MacInnis also notes that OpenAI's two effort settings of the same model landed on either side of Opus 4.6, so the setting mattered more than the brand [16].

Newer was not better in the one case he retested. Re-running the same 2,100 attempts on Grok 4.6 dropped accuracy from 87.3% to 85.9% and nearly doubled typical response time, from 71 seconds to 131 [17], roughly 1.8 times slower [5]. MacInnis flags his own caveat: that 87.3% does not match the 89.1% in his published table because it came from a different run on a different day, and only results measured in the same run should be compared [18].

The other number to internalise is the tuning. Rippling spent five months tuning instructions and tools specifically around Opus 4.6, worth a point or two by MacInnis's own estimate [19][20]. That is the same order as the gap between the leader and the runner-up, which makes the leaderboard position a property of the harness as much as the model.

Watch what your own harness does when you swap models rather than what a release note claims, and treat a new version as a reason to re-run the suite, not to upgrade [21]. Watch, too, whether any other operator publishes graded production data at this size; for now this is one company's test on one product surface, and the numbers are only portable to work that is similarly constrained.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories