Invest1 publisher3 min readPublished
Larridin's benchmark puts the median engineer's AI coding bill at $213 a week
Larridin says the median engineer draws $213 a week in billed AI-coding spend, with the 90th percentile at $911. Its cohort data finds extra spend lifts output only for engineers already fluent with the tools, so each team must measure its own return before raising budgets.
The Investor · Invest desk

What happened
- Larridin built the benchmark from production billing and engineering telemetry on engineers who merged code and drew billed AI-coding spend in the four weeks to August 2, 2026.
- Output is Larridin's own score, with each merged pull request rated by a model on five complexity levels, discounted for weak quality and missing tests, and scaled by code churn.
- The most AI-native engineers, with 79% of output AI-attributed, reached 11.8x output at about $1,300 a week and shipped roughly twice what partial adopters did at equal spend.
- Engineers with 15% or less of their output AI-attributed topped out near 1.9x and stayed flat as weekly spend ranged from $21 to $421.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- decision Budget owners have to set review triggers team by team from their own output curves, because the data shows no single token cap suits a fluent team and a low-adoption team at once.
- contradiction A CFO who sizes a cap off the tenfold spread is using a gap the two published percentiles do not reproduce, so any cap built on it rests on a figure the write-up leaves unshown.
- constraint Because the results are associational, the benchmark cannot yet tell a finance team whether a bigger budget would move any engineer up the curve, only which engineers already sit high on it.
Annualised, $213 a week comes to about $11,100 per engineer, and $911 comes to about $47,400 [1][2]. A 100-engineer team at the median spends roughly $1.1 million a year on metered AI-coding tokens [4]. Larridin calls its figures a floor, since tokens consumed on flat-fee plans never reach a metered bill [4]. SaaStr's write-up adds that the spend is spread across provider invoices, corporate cards and personal subscriptions, where most CFOs cannot see it [5].
The tenfold spread does not follow from the two published percentiles. $911 is 4.3 times $213 [3], so the benchmark's "more than 10x" has to compare the 90th percentile with some lower one [2]. The write-up does not print that lower figure. The gap a finance team can check from what was published runs from the median engineer to the one at the 90th percentile, and it is 4.3 times [3].
The cohort results support three readings, and they send budgets in different directions. One: spend buys output once engineers know the tools. The money then goes to the fluent, and partial adopters get a review trigger near $600 a week, where their marginal payoff fell by about half [9], or roughly $31,200 a year per engineer [6]. Or output drives spend. Larridin allows for this, noting that high-output engineers may spend more because they ship more [13], and if so a bigger budget moves nobody up the curve. A third possibility is billing mix: an engineer who works mostly on a flat-fee plan looks cheap on the metered bill [4], so some of the spread may be plan choice.
In my view the low-AI result holds under all three readings [6]. "Extra dollars bought activity, and no additional output," SaaStr wrote [7]. The $400 a week between the bottom and top of that cohort's spending range is about $20,800 a year per engineer with no measured return [5]. Even if some low spenders hide flat-fee usage, the ones at the top of the range were metered and still flat. For that group I'd put the next dollar into training before raising anyone's token cap. Larridin's own advice is to track each team's curve and set a review trigger where it levels off [11].
The fluent-versus-partial comparison, the strongest argument for tying budgets to skill, comes from one company, with the same tools, the same prices and a common starting spend of about $170 a week [10]. If the curves fail to repeat at other companies, or engineers given bigger budgets stay at the same output, the case for fluency-based budgets goes with them. The data also comes from a vendor selling the measurement. SaaStr featured Larridin as its AI App of the Week and called the benchmark "the best argument for the product" [14][15].
What to watch
- Whether Larridin publishes the lower percentile behind its 'more than 10x' spread, to show how much of the gap sits below the median.
- A second benchmark drawing on more than one company, testing whether fluent engineers still ship about twice as much as partial adopters at equal spend.
- Whether Larridin finds a way to count flat-fee plan usage, moving its per-engineer figures off a floor.