Invest1 publisher3 min readPublished
Rippling graded 2,100 agent runs per model. The cheap one basically tied the flagship.
A production test across 15 models put seven of them inside a one-point spread on pass rate. On constrained payroll work, the price premium bought speed, not correctness.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Rippling's President and CPO Matt MacInnis published a test of AI models doing real work inside a real production system, rather than a leaderboard or invented tasks.
- The study covered about 2,100 graded attempts per model, across 15 models, on actual personnel, payroll, and financial records.
- Seven models landed between 88.5% and 89.5% pass rate, a one-point spread; adding the leader at 91.0% makes the whole competitive group 2.5 points wide.
- MacInnis writes that defaulting to whatever the big lab just shipped means paying a 3x to 7x premium for a difference customers cannot detect.
- Every attempt either passed Rippling's production correctness checks or failed, with no partial credit; an attempt that never finished counted as a failure.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Rippling's president and chief product officer, Matt MacInnis, published what most B2B companies keep internal: roughly 2,100 graded attempts per model, across 15 models, run against actual personnel, payroll and financial records [1][2]. The consequence for anyone signing an inference bill is that seven models finished between 88.5% and 89.5% pass rate, a one-point spread, while the prices attached to that band differ by up to seven times [3][4].
The grading is why the result carries weight. Every attempt either passed Rippling's production correctness checks or failed, with no partial credit, and an attempt that never finished counted as a failure [5]. According to MacInnis, that is a harder grader than most published benchmarks use [6]. The work was narrow and operational: headcount by department, tenure distribution, raising base salaries for everyone meeting a condition by 10%, onboarding a hire through a checklist, scheduling a termination and routing its approvals, entering amounts into a pay run from a spreadsheet [7].
On those tasks, 15 models collapse to three defensible purchases. Opus 4.6 leads accuracy at 91.0% for $1,453 per full test run, with 154 seconds at the slowest 10% [8]. GPT-5.5 at medium effort scored 89.5% for $1,435, an $18 gap [9] that is about 1.2% of the bill [1]. GPT-5.5 at low effort scored 88.8% for $1,308 with a 130-second tail [10]. GLM 5.2 scored 88.7% for $621 with a 243-second tail [11]: 57% cheaper than the accuracy leader for 2.3 points of pass rate [2], and 53% cheaper than GPT-5.5 low for a tenth of a point, at the cost of 89 extra seconds at the tail versus Opus 4.6 [3][4].
The expensive end does not survive that comparison. Fable 5 was the most expensive model tested and finished fifth, matching GPT-5.5 low on accuracy at 3.3 times the price and twice the wait [12][13]. Opus 5 was beaten on both price and speed by Grok 4.5, which scored 0.2 points behind it for 68% less [14]. Opus 4.6 beat both newer Anthropic models on accuracy while costing 42% less than Opus 5 and a third of Fable 5 [15]. MacInnis also notes that OpenAI's two effort settings of the same model landed on either side of Opus 4.6, so the setting mattered more than the brand [16].
Newer was not better in the one case he retested. Re-running the same 2,100 attempts on Grok 4.6 dropped accuracy from 87.3% to 85.9% and nearly doubled typical response time, from 71 seconds to 131 [17], roughly 1.8 times slower [5]. MacInnis flags his own caveat: that 87.3% does not match the 89.1% in his published table because it came from a different run on a different day, and only results measured in the same run should be compared [18].
The other number to internalise is the tuning. Rippling spent five months tuning instructions and tools specifically around Opus 4.6, worth a point or two by MacInnis's own estimate [19][20]. That is the same order as the gap between the leader and the runner-up, which makes the leaderboard position a property of the harness as much as the model.
Watch what your own harness does when you swap models rather than what a release note claims, and treat a new version as a reason to re-run the suite, not to upgrade [21]. Watch, too, whether any other operator publishes graded production data at this size; for now this is one company's test on one product surface, and the numbers are only portable to work that is similarly constrained.