Skip to content

Build1 publisher3 min readPublished

Four months of A100 bills say self-hosting is a utilization bet, not a cost saving

One engineer's home lab tally: $1,400 a month for two A100s running 40% idle, against open-weight models he measured inside noise of GPT-4o. The break-even is real, and it sits high.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The author ran a four-month home lab experiment asking whether self-hosting open-source models still makes economic sense versus API access.
  • The test matrix covered 3 distinct traffic tiers, 10 model candidates and 2 hosting modes, which the author says gave him enough confidence to stop renting his A100.
  • Six months ago the author was spending roughly $1,400 per month on two A100 80GB instances running a 32B parameter model.
  • Roughly 40% of the author's GPU hours were idle.
  • The author says the DevOps overhead of self-hosting was eating his weekends.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An engineer running a home lab published four months of measured spend comparing self-hosted open-weight models against hosted API endpoints, and stopped renting his A100 capacity at the end of it [1][2]. The useful part is not the verdict but the utilization figure underneath it: roughly 40 percent of his GPU hours were idle on a bill of about $1,400 a month for two A100 80GB instances serving a 32B parameter model [3][4]. That is about $560 a month of capacity bought and not used [1]. For scale, his mid-tier traffic scenario, 1.5 billion tokens a month or 50 million a day, costs $375 a month in output tokens at the $0.25 per million rate he benchmarked against [14][16][3]. The idle share of one GPU bill was worth more than the entire API bill for a workload most teams would call serious [5]. The second number is the one people skip. According to the author, GPU rental was only about 55 percent of his actual cost over six months [11]. Applied to that $1,400 line, it implies a true monthly cost nearer $2,545, with roughly $1,145 in everything else [2]. He reports a $3,050 total sitting inside a $900 to $4,900 monthly hidden-cost band, and says the relationship between models served concurrently and DevOps hours was nearly linear in his logs [12][13]. The weekend cost was, in his telling, part of what drove the decision [5]. On quality, he says open-weight models landed within statistical noise of GPT-4o on summarization and code generation across 500 prompts from his own dataset, consistent over three separate runs [6]. Take that for what it is: one domain, one eval set, self-reported. It removes capability as an excuse for renting GPUs in his workload. It does not settle the general case. The three tiers are where the argument holds up. At 1 million tokens a day, output cost is $7.50, or about $12.50 once you apply his measured 3:1 input-to-output cost ratio, against a $400 to $800 floor for the smallest viable GPU even when idle [15]. That is the 32x gap he claims, and it checks out arithmetically [6]. At 50 million a day, API runs 3 to 5 times cheaper, with on-prem hardware amortized over 36 months landing at $500 to $1,000 [16]. At the top tier, DeepSeek V4 Flash costs $3,750 a month in output alone, Qwen3-32B $4,200, cloud rental $4,000 to $8,000, and amortized on-prem $2,000 to $4,000, which he calls a tie [17]. He also concedes the standard deviation in those projections is larger than the mean difference [18]. Which is where the headline conclusion gets loose. He closes by saying API access is cheaper until roughly 50 million tokens a day [19]. But the tier where his own costs actually tie implies 15 billion output tokens a month, about 500 million a day [4]. His numbers put the crossover an order of magnitude above his stated threshold, and even there self-hosting only edges ahead if the hardware is already bought and someone is already paid to operate it [17]. Two notes on the pricing. All ten price points came from a single aggregator, Global API, which is also the product named in the takeaway [7][19]. And the price distribution is ten rows: median $0.345 per million output, a mean pulled up by Hunyuan-A13B and GLM-4-32B, and three models between $0.19 and $0.28 [8]. Three prices spread across a range are a cluster, not a mode. The $0.01 per million rates on Qwen3-8B and GLM-4-9B are real for prototyping and irrelevant to capacity planning [9]. Self-host ranges stay wide because they depend on whether you take Lambda Labs spot pricing or amortize your own boxes [10]. The number to instrument before running this comparison yourself is idle percentage, because it is what converts a fixed GPU commitment into an effective per-token price.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories