Skip to content

Build1 publisher3 min readPublished

A $3,499 Mac Studio saves 22 cents a day against hosted inference in Sunk Cost's model

The calculator prices its featured coding workload at about $6.94 a month through OpenRouter against 26 cents of electricity. The machine is still $3,419 down after a year and 7.97 billion tokens from break-even.

The Engineer · Build desk

Illustration accompanying A $3,499 Mac Studio saves 22 cents a day against hosted inference in Sunk Cost's model

What happened

  • Sunk Cost's featured example pairs a $3,499 M5 Max Mac Studio and 64 GB of unified memory with Alibaba's quantized Qwen3.8 27B model, running 500,000 tokens a day at 15 input tokens per output token.
  • The calculator prices that workload at roughly $6.94 a month through an OpenRouter endpoint against 26 cents of monthly electricity, using rates of 32 cents per million input tokens and $2.50 per million output.
  • At 22 cents of daily saving the machine is $3,419 underwater after its first year and reaches break-even after about 7.97 billion tokens, which Sunk Cost puts 44 years out.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A buyer at this usage level has to justify the $3,499 on privacy, offline operation, outage tolerance and resale, four things Sunk Cost leaves out of the sum on purpose, because $80 a year of token saving will not carry the purchase.
  • constraint Latency argues against the local box in this configuration: the operator absorbs roughly 91 hours of extra waiting a year for that $80, which is 88 cents an hour of saving.
  • capability Anyone weighing a local machine can now put their own prompt-to-output ratio, context size and expected annual API decline into the sum before ordering hardware.

Two inputs produce the 44 years. Sunk Cost prices the featured workload at about $6.94 a month through an OpenRouter endpoint and 26 cents a month in electricity [6]. The gross saving is $6.68 a month, roughly $80 a year, and $3,499 divided by $80 is 43.6 years [1][2].

Most of that hosted bill is prompt. At 468,750 input tokens a day and 32 cents per million, input costs about $4.56 a month; the 31,250 daily output tokens cost about $2.38, so two thirds of the bill is input [5][7][4]. Sunk Cost does not price prompt caching, and prompt caching works on the input side [16].

The two assumptions most likely to be wrong move the answer very little. Generation speed is an estimate of 24.7 tokens per second, derived from memory bandwidth, an assumed bytes-per-token figure and a 75% efficiency factor, on a configuration nobody has measured in production [10][11]. Power is 145 watts under load, taken from Apple's published maximum for the previous M4 Max Mac Studio [12]. Delete the electricity line entirely and break-even goes from 44 years to 42 [5].

Working backwards from the published inputs, 614 GB/s at 75% efficiency divided by 24.7 tokens per second implies about 18.6 GB read per token [9][6]. In a dense model the weights stream from memory once per token, so the rate is just a bandwidth quotient. Sunk Cost labels it an estimate [10].

The speed estimate matters more on the waiting side. Sunk Cost puts a 1,000-token response at 40 seconds on the Mac and 13 seconds through an API running at 80 tokens per second, and totals the extra delay at about 15 minutes a day [15]. That checks out: 31,250 output tokens is about 31 responses of 1,000 tokens, 27 seconds slower each, or 14 minutes [7]. Over a year the operator waits about 91 hours to save $80, which is 88 cents an hour [8].

For payback inside three years at these prices, the workload has to carry about $97 a month of hosted spend. That is roughly 15 times the modeled one: 7.3 million tokens a day, and about five hours a day of local generation at 25 tokens per second [9]. It also assumes the September 3rd rates hold [7]. Sunk Cost includes an annual API decline input, and RuntimeWire's reason for it is that hosted vendors spread new hardware and utilization gains across many customers while the Mac owner captures none of those improvements unless faster software or a better local model arrives [14][18].

Sunk Cost excludes the value of keeping code and documents on-device, working offline, riding out provider outages, using the Mac for unrelated work, and resale [16]. Those are the arguments a buyer has left. The calculator's own numbers give the API 80 tokens per second against the Mac's 25, so latency is not one of them for this configuration [15][10].

What to watch

  • Measured generation rates on shipping M5 Max hardware after the September 22nd availability date, against the calculator's 24.7 tokens per second.
  • Any cut to the 32 cents per million input rate: each one lengthens the payback beyond 44 years.
  • Faster local software or a better model that fits 64 GB, the one route RuntimeWire identifies that improves the owner's side without new hardware.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories