Skip to content

Build1 publisher3 min readPublished

LessWrong cost study prices the median hour of AI task work at 4 cents of GPU time

LessWrong post puts the GPU cost of an AI doing an hour of median human work at about 4 cents, against a $25 US median wage. Current API prices narrow that gap sharply, and for the hardest tasks they lift AI cost to the hourly rate of a skilled engineer.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying LessWrong cost study prices the median hour of AI task work at 4 cents of GPU time
Photo: lesswrong.com

What happened

  • Even the expensive tasks came to about $15 per hour of human work when priced at H100 hardware cost.
  • Where the cost of an AI run was known, the price paid per FLOP came out more than ten times the hardware cost.
  • Frontier labs do not publish compute, parameter counts or architecture, so the post's compute figures are usually guesses.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Priced at a tenfold markup, the median task still comes in about 60 times under a $25 wage, so routine agent work keeps its budget case even at invoice rates.
  • cost Teams automating the hardest tasks pay at least $150 per hour of human-equivalent work at current prices, so any saving depends on what their own engineers cost fully loaded.
  • constraint Guessed compute and a cheapest-successful-run filter keep the figures at order-of-magnitude, and they apply only to teams whose runs match the dataset's best case.
  • precedent If serving prices fall toward cost per FLOP as the author expects, the hard-task figure drops back toward $15 an hour and the budget case extends to the expensive tail.

The method multiplies two estimates. One is the compute a model used to finish a task. The other is how long a human would take to do the same task [5]. The ratio, FLOP per second of human time, is then priced at what a modern H100 costs to run [15]. The author wrote that both estimates are crude [6]. Frontier labs do not publish compute, parameter counts or architecture, so the compute side is usually a guess, and human task time varies [6]. The database was compiled with GPT-6 Astra and Fable 5.1 [5], so the models helped draw up their own cost sheet. The post claims order-of-magnitude accuracy and no more [16].

The most useful result is the shape of the data. Compute rose roughly linearly with human time [8]. Scatter is substantial, but outside a cluster of very cheap tasks there is no obvious trend in FLOP per second of human time [8]. A flat per-hour rate means agent cost grows with task length the way payroll does. The costliest points are a Navier-Stokes proof, a Pokemon Crystal play-through and an ARC-AGI-3 environment [9].

Each plotted point is the cheapest system that reached at least human performance on that task [7]. More expensive runs of the same tasks appear only in the interactive version [7]. The 4-cent median is a best case per task. It transfers to a team's own work only if that team's agent reaches the cheapest successful configuration and is billed at hardware cost.

Billing is the larger gap. Where the post knew what a run actually cost, the price per FLOP was more than ten times the hardware figure [10]. The author attributes the markup to labs covering training and research costs and taking large margins [11]. At that floor the median task costs at least 40 cents per hour of human work [1]. At exactly ten times, that is about 60 times below the $25 US median wage [3]. The expensive tasks cost at least $150 an hour [2]. That sits inside the $100 to $300 fully loaded hourly cost the post gives for research scientists and software engineers [4].

The author argues that in the long run the marginal cost of serving a model is set by cost per FLOP, as training and research spending is amortized over more use [12]. That is a forecast about pricing. If it holds, API prices converge on the hardware figures. "Doing a task is not the same as doing it cheaply," the author wrote [13].

The post concludes that "when AI systems can do a human task, they almost always are far cheaper than humans" [14]. I think the evidence supports that for the median task at every price the post reports. For the hardest tasks it holds at hardware cost. At current prices those tasks cost at least as much as the low end of a skilled engineer's loaded rate [2][4].

What to watch

  • Frontier labs publishing per-task compute or run costs would replace the guessed FLOP figures behind every point in the dataset.
  • API price per FLOP falling toward hardware cost, as the author's amortization argument expects; at a tenfold markup the hardest tasks sit at engineer rates.
  • The interactive version's more expensive runs, if their per-hour costs sit far above the cheapest successful run on the same tasks.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories