Build1 publisher3 min readPublished
One DGX Spark, four Macs, and an attempt to turn token billing into a capital purchase
A developer moved both routing decisions and code generation onto local models after round-the-clock agents made usage billing permanent. The CI arithmetic is published. The hardware price is not.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- From late June into July 2026 the author rebuilt the execution backbone of their development onto their own local large language models; the trigger was cost, driven by 24/7 operation sending every routing decision and every code generation to a cloud AI with usage-based billing that recurs monthly.
- The author bought one NVIDIA DGX Spark and combined it with four Macs already owned to build an execution backbone that development tasks flow through.
- The author moved the task-routing decisions (called the orchestrator) and much of the hands-on work from the cloud AI to local LLMs.
- The stated hypothesis was that buying hardware once and shifting execution onto local LLMs could erase most of the ongoing cost, replacing usage-based billing that grows with use with a one-time hardware cost.
- Leaving the hands-on work to the AI and keeping only decisions for the human, pushed far enough, is described as Human-Out-Of-The-Loop (HOOTL), where the human steps outside the loop.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Between late June and July 2026 a developer rebuilt the execution backbone of their own development work onto local large language models, buying one NVIDIA DGX Spark and combining it with four Macs already on hand [1][2]. The stated trigger was cost: hand development to agents that run 24 hours a day and every task-routing decision plus every code generation goes to a cloud model, so usage billing accrues in proportion to use, every month, indefinitely [1][3].
The thesis is straightforward: buy the hardware once, shift execution onto it, and convert growing usage-based billing into a one-time cost [4]. Two reasons are given, and the second matters more than the first. One is the recurring bill. The other is not wanting the decision-making itself to depend on external billing [12]. The routing layer, which the author calls the orchestrator and names local-commander, runs continuously, so no cloud model is placed on it as a design principle [13]. If the external service stops, development stops with it [14]. Cloud spend halts work in two shapes: a free tier that hits its ceiling, or metered billing that piles up without limit [15]. That is not hypothetical for this team. Their organisation's GitHub Actions stopped from April 2026, on suspicion of hitting the free-tier ceiling [16].
The only arithmetic actually published concerns CI, not the inference hardware. Two machines' worth of GitHub Actions standard runners (Linux, 2 cores, $0.008/min), at an estimated half utilisation of 12 hours a day, comes to roughly JPY 55,000 a month; two self-hosted CI servers cost JPY 130,000 each, JPY 260,000 once, which the author puts at about five months to break even [17][18]. Dividing those figures gives 4.7 months, so the estimate is sound [20]. The dollar side works out to about $345.60 a month, implying an exchange rate near JPY 159 to the dollar, which is worth knowing before you reuse the number [21]. No price is quoted for the DGX Spark or for the total hardware outlay [24], and the author says plainly that this is not a claim about a specific saving but about how to wire the recurring cost out where the design allows [6].
The purchase gate was measurement, locked in an architecture decision record before buying so that wanting the machine could not decide it [7]. The measured baseline explains why: a 31B-class model on a MacBook Pro (M3) delivered an effective 5 tok/s and 60 to 200 seconds per decision, which the author judged too slow for continuous operation [19]. At those latencies a single orchestrator lane tops out somewhere between 432 and 1,440 decisions a day [22], on budgets of roughly 300 to 1,000 tokens each [23]. Model placement was then decided per role from measured speed and cost: routing on a 14B model, hands-on code generation on what is described as a zero-billing 72B lane [8]. One model per host, to avoid swap costs and to keep the local-versus-cloud split visible in numbers [9]. The stated end state is Human-Out-Of-The-Loop, with the human holding decisions and nothing else [5].
Three things to watch. The DGX Spark arrived on 10 July, took about two weeks to reach a working form, and then an incident occurs in late July, which is where the useful engineering usually lives [11]. The material does not specify how the 72B lane achieves zero billing [25], and that is the load-bearing assumption in the whole scheme. And capex conversion relocates cost rather than deleting it: power draw, depreciation and the hours spent lining up four Macs and a Spark stay on the books whether or not anyone bills you for them.