Invest1 distinct publisher3 min readUpdated
A production test across 15 models put seven of them inside a one-point spread on pass rate. On constrained payroll work, the price premium bought speed, not correctness.
The Investor · Invest desk
Compiled by The InvestorSomething wrong?How this is made
Rippling's president and chief product officer, Matt MacInnis, published what most B2B companies keep internal: roughly 2,100 graded attempts per model, across 15 models, run against actual personnel, payroll and financial records [1][2]. The consequence for anyone signing an inference bill is that seven models finished between 88.5% and 89.5% pass rate, a one-point spread, while the prices attached to that band differ by up to seven times [3][4].
The grading is why the result carries weight. Every attempt either passed Rippling's production correctness checks or failed, with no partial credit, and an attempt that never finished counted as a failure [5]. According to MacInnis, that is a harder grader than most published benchmarks use [6]. The work was narrow and operational: headcount by department, tenure distribution, raising base salaries for everyone meeting a condition by 10%, onboarding a hire through a checklist, scheduling a termination and routing its approvals, entering amounts into a pay run from a spreadsheet [7].
On those tasks, 15 models collapse to three defensible purchases. Opus 4.6 leads accuracy at 91.0% for $1,453 per full test run, with 154 seconds at the slowest 10% [8]. GPT-5.5 at medium effort scored 89.5% for $1,435, an $18 gap [9] that is about 1.2% of the bill [1]. GPT-5.5 at low effort scored 88.8% for $1,308 with a 130-second tail [10]. GLM 5.2 scored 88.7% for $621 with a 243-second tail [11]: 57% cheaper than the accuracy leader for 2.3 points of pass rate [2], and 53% cheaper than GPT-5.5 low for a tenth of a point, at the cost of 89 extra seconds at the tail versus Opus 4.6 [3][4].
The expensive end does not survive that comparison. Fable 5 was the most expensive model tested and finished fifth, matching GPT-5.5 low on accuracy at 3.3 times the price and twice the wait [12][13]. Opus 5 was beaten on both price and speed by Grok 4.5, which scored 0.2 points behind it for 68% less [14]. Opus 4.6 beat both newer Anthropic models on accuracy while costing 42% less than Opus 5 and a third of Fable 5 [15]. MacInnis also notes that OpenAI's two effort settings of the same model landed on either side of Opus 4.6, so the setting mattered more than the brand [16].
Newer was not better in the one case he retested. Re-running the same 2,100 attempts on Grok 4.6 dropped accuracy from 87.3% to 85.9% and nearly doubled typical response time, from 71 seconds to 131 [17], roughly 1.8 times slower [5]. MacInnis flags his own caveat: that 87.3% does not match the 89.1% in his published table because it came from a different run on a different day, and only results measured in the same run should be compared [18].
The other number to internalise is the tuning. Rippling spent five months tuning instructions and tools specifically around Opus 4.6, worth a point or two by MacInnis's own estimate [19][20]. That is the same order as the gap between the leader and the runner-up, which makes the leaderboard position a property of the harness as much as the model.
Watch what your own harness does when you swap models rather than what a release note claims, and treat a new version as a reason to re-run the suite, not to upgrade [21]. Watch, too, whether any other operator publishes graded production data at this size; for now this is one company's test on one product surface, and the numbers are only portable to work that is similarly constrained.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Rippling's President and CPO Matt MacInnis published a test of AI models doing real work inside a real production system, rather than a leaderboard or invented tasks.
The study covered about 2,100 graded attempts per model, across 15 models, on actual personnel, payroll, and financial records.
Every attempt either passed Rippling's production correctness checks or failed, with no partial credit; an attempt that never finished counted as a failure.
MacInnis describes the grader as much harder than most published benchmarks use.
Tasks included questions such as headcount by department and tenure distribution, and actions such as increasing base salaries of everyone meeting a condition by 10%, onboarding a new hire through the checklist, scheduling a termination and routing its approvals, and entering payment amounts into a pay run from a spreadsheet.
Rippling spent five months tuning its instructions and tools specifically around Opus 4.6.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Unusually concrete methodology, entirely unreplicated
The underlying test is far better specified than a typical vendor claim: ~2,100 graded attempts per model across 15 models, binary pass/fail against production correctness checks with timeouts scored as failures, and three reported metrics per model including tail latency. That raises evidence quality well above assertion. It is held down by everything the supplied material lacks: a single publisher retelling a single self-published study, no link to or inspection of the raw data, grader definition or prompt harness, no computed confidence intervals behind the 'inside the margin of error' language, an acknowledged internal inconsistency between the 87.3% and 89.1% Grok 4.5 figures, and an asymmetric design in which only the winning model received five months of tuning.
One production deployment, one company, one domain
There is real deployed usage rather than a demo: Rippling discloses a live single-model, two-tool agent tuned for five months against Opus 4.6, operating on actual payroll, onboarding and termination workflows, plus two benchmark runs including a re-test on a newer model. But adoption breadth is minimal in the supplied material — one company, one task domain, no other named adopters, no disclosed traffic volume, seat counts or customer-facing rollout, and no evidence that anyone acted on the study's cheap-model recommendation.
Headline framing outruns the measured deltas
The 'cheapest tied the most expensive' framing is stronger than the numbers support: GLM 5.2 is 2.3 points below the leader and 89 seconds slower at the slowest decile, which is a real trade rather than a tie, and the 3x-7x premium argument leans on a single one-tenth-of-a-point pairing. The Grok 4.6 regression is presented as proof that newer is worse while resting on a baseline the article itself says is not run-comparable. Offsetting the overreach, the source volunteers its own caveats — margin-of-error language, untuned competitors, the mismatched baseline — and the core finding of a compressed 2.5-point quality band with wide price dispersion is directly measured, so the gap is modest rather than large.
Vendor-executive study, founder-audience amplifier
The study is authored by Rippling's own President and CPO, and its two flattering conclusions — that five months of in-house tuning is worth a whole model generation and that the sensible move is to stay on the incumbent configuration — happen to validate Rippling's own engineering investment while positioning the company as an unusually rigorous AI operator. The publisher packages it as seven numbered takeaways for B2B founders, a format aligned with its audience-growth interest. Score is mid-range rather than high because the supplied material shows no vendor sponsorship, no model provider paying for placement, and the author publishes results unfavourable to several vendors including the one he tuned toward, plus self-critical caveats.
Directionally credible, structurally single-sourced
Confidence is moderate. The claim set is internally detailed and the source is candid about its own limits, which supports the directional conclusion that top-end quality has compressed while price has not. But every number traces to one publisher's account of one vendor executive's self-published test, the design is asymmetric in the leader's favour, one headline comparison uses acknowledged non-comparable runs, and the findings are scoped to constrained payroll and HR tasks with no basis in the supplied material for generalizing further.
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
invest
Canva guided 2026 growth down to 20%. The worse detail is who never got asked.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.