Build1 distinct publisher3 min readUpdated
One engineer's home lab tally: $1,400 a month for two A100s running 40% idle, against open-weight models he measured inside noise of GPT-4o. The break-even is real, and it sits high.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
An engineer running a home lab published four months of measured spend comparing self-hosted open-weight models against hosted API endpoints, and stopped renting his A100 capacity at the end of it [1][2]. The useful part is not the verdict but the utilization figure underneath it: roughly 40 percent of his GPU hours were idle on a bill of about $1,400 a month for two A100 80GB instances serving a 32B parameter model [3][4]. That is about $560 a month of capacity bought and not used [1]. For scale, his mid-tier traffic scenario, 1.5 billion tokens a month or 50 million a day, costs $375 a month in output tokens at the $0.25 per million rate he benchmarked against [14][16][3]. The idle share of one GPU bill was worth more than the entire API bill for a workload most teams would call serious [5]. The second number is the one people skip. According to the author, GPU rental was only about 55 percent of his actual cost over six months [11]. Applied to that $1,400 line, it implies a true monthly cost nearer $2,545, with roughly $1,145 in everything else [2]. He reports a $3,050 total sitting inside a $900 to $4,900 monthly hidden-cost band, and says the relationship between models served concurrently and DevOps hours was nearly linear in his logs [12][13]. The weekend cost was, in his telling, part of what drove the decision [5]. On quality, he says open-weight models landed within statistical noise of GPT-4o on summarization and code generation across 500 prompts from his own dataset, consistent over three separate runs [6]. Take that for what it is: one domain, one eval set, self-reported. It removes capability as an excuse for renting GPUs in his workload. It does not settle the general case. The three tiers are where the argument holds up. At 1 million tokens a day, output cost is $7.50, or about $12.50 once you apply his measured 3:1 input-to-output cost ratio, against a $400 to $800 floor for the smallest viable GPU even when idle [15]. That is the 32x gap he claims, and it checks out arithmetically [6]. At 50 million a day, API runs 3 to 5 times cheaper, with on-prem hardware amortized over 36 months landing at $500 to $1,000 [16]. At the top tier, DeepSeek V4 Flash costs $3,750 a month in output alone, Qwen3-32B $4,200, cloud rental $4,000 to $8,000, and amortized on-prem $2,000 to $4,000, which he calls a tie [17]. He also concedes the standard deviation in those projections is larger than the mean difference [18]. Which is where the headline conclusion gets loose. He closes by saying API access is cheaper until roughly 50 million tokens a day [19]. But the tier where his own costs actually tie implies 15 billion output tokens a month, about 500 million a day [4]. His numbers put the crossover an order of magnitude above his stated threshold, and even there self-hosting only edges ahead if the hardware is already bought and someone is already paid to operate it [17]. Two notes on the pricing. All ten price points came from a single aggregator, Global API, which is also the product named in the takeaway [7][19]. And the price distribution is ten rows: median $0.345 per million output, a mean pulled up by Hunyuan-A13B and GLM-4-32B, and three models between $0.19 and $0.28 [8]. Three prices spread across a range are a cluster, not a mode. The $0.01 per million rates on Qwen3-8B and GLM-4-9B are real for prototyping and irrelevant to capacity planning [9]. Self-host ranges stay wide because they depend on whether you take Lambda Labs spot pricing or amortize your own boxes [10]. The number to instrument before running this comparison yourself is idle percentage, because it is what converts a fixed GPU commitment into an effective per-token price.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author ran a four-month home lab experiment asking whether self-hosting open-source models still makes economic sense versus API access.
The test matrix covered 3 distinct traffic tiers, 10 model candidates and 2 hosting modes, which the author says gave him enough confidence to stop renting his A100.
Six months ago the author was spending roughly $1,400 per month on two A100 80GB instances running a 32B parameter model.
The author narrowed the experiment to ten models with publicly verifiable pricing via Global API, focusing on output-side rates because that is where cost compounds.
The median API price across the ten models is $0.345 per million output tokens; the mean is dragged up by Hunyuan-A13B and GLM-4-32B; three models cluster around $0.19 to $0.28.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One unaudited home-lab ledger
All evidence is a single self-published first-person post. Cost figures, idle percentage, hidden-cost share and eval results are self-reported with no invoices, logs, dataset or scores published, and no second source corroborates any number. The internal arithmetic that can be checked does hold - the low-tier 32x gap and the 1.5B/month to 50M/day conversion both reproduce - but one derived line conflicts with the stated takeaway, and the top-tier comparison is admitted to have a projection spread wider than the difference it measures.
A single practitioner's migration
The only observable adoption is one anonymous engineer moving his own home-lab workload from rented A100s to hosted open-weight APIs, plus his private benchmark run and a snapshot of one provider's price list. There is no organizational deployment, no user or revenue disclosure, no third-party usage data, and no evidence that anyone else has followed the same path.
Verdict framing outruns one home lab
The piece is more candid than most - it labels the top tier a tie and warns against 'API is always cheaper' - but the packaging still overreaches its evidence. A headline 'verdict', a named ~50M tokens/day crossover and a GPT-4o parity finding are generalized from one engineer's private dataset and one vendor's price list, and the stated crossover conflicts with the article's own top-tier arithmetic by roughly ten times. The durable, well-supported part is narrower than advertised: at his utilization, self-hosting was a utilization and staffing bet.
Single-vendor pricing and embedded endpoint
Every price in the analysis comes from one catalog, Global API; the recommendation names that provider by brand; and the article ships working client code hard-coding https://global-apis.com/v1 with a GLOBAL_API_KEY environment variable. That structure matches vendor-adjacent developer content, and no relationship, affiliate arrangement or sponsorship is disclosed either way. The counterweight is that the author also publishes findings unfavorable to the simple 'API always wins' pitch, including a tied top tier.
Low - one voice, checkable in parts
Confidence is limited by a single-publisher, single-author cluster with no corroboration and undisclosed vendor proximity. It is not lower because the reasoning is transparent enough to audit: several figures reproduce arithmetically, the author states his own uncertainty at the top tier, and the one internal contradiction is identifiable from the text itself. Directional read on utilization economics is reasonable; specific dollar thresholds should not be relied on.
build
Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k1 distinct publisher
leadership
Re-baseline AI procurement on cost per completed task, not dollars per million tokens1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 16, 2026