Build1 distinct publisher3 min readUpdated
Its August 20 bake-off priced three agents on one Nemotron port. The claim that findings transfer to the next model-chip pair is the one with no number attached.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The load-bearing sentence in the August 20 write-up is not in the cost table. It is the statement that Base Compute is trying to preserve what its agents learn on one optimization job and apply those findings to later models and processors [4]. That is the difference between a consulting practice and a product. It is also the part the published experiment cannot speak to, because all three agents were handed the same already-completed Nemotron port, the same evaluation setup and the same eight-hour limit [5]. What was measured is how cheaply a model searches a fixed problem. Whether anything it finds survives contact with the next processor is asserted, not shown.
The reason that labor exists at all is architectural. BaseRT, introduced on July 1, is a native C++ runtime that goes at Apple's Metal API directly with handwritten kernels, fused operators and hardware-specific dispatch rather than routing through MLX, PyTorch or Core ML [11]. Declining the framework layer buys control and hands you the bill: every new architecture can otherwise mean another round of operator implementation, numerical verification and device-specific optimization by systems engineers [3]. The first public target, NVIDIA's Nemotron 3 Nano 30B-A3B at 52 layers, was a model BaseRT did not support [13]. Agentic kernel work is the repayment plan for the framework Base Compute chose not to use.
Now the arithmetic the write-up leaves on the floor. Claude Fable 5 took the top score with 50 experiments at $265; Kimi K3 reached 97 percent of it in 21 experiments for $77 [6][7]. So the final three percentage points cost $188, about $63 a point [15]. Per experiment, Fable ran at $5.30 against Kimi's $3.67, roughly 44 percent more for each attempt, and it made 2.4 times as many attempts [16][18]. Line the three runs up and the curve is steep then flat: going from 12 experiments to 21 bought 11 points, and the next 29 experiments bought 3 [17]. Any budget set by "run it to the eight-hour limit" is mostly paying for search that repeats itself.
GLM 5.2 is the awkward row. It reached 86 percent in 12 experiments through BaseRT on an M3 Ultra Mac Studio, and Base Compute assigned that run near-zero marginal cost [8]. "Marginal" is doing heavy lifting there: the Mac and the eight hours it was occupied are capital and opportunity, not zero.
One detail cuts in Base Compute's favor. All three agents flagged early changes involving state-space-model prefill, decode dispatch and the Mamba convolution path before their results diverged [9]. If the profitable moves are that legible, they look like properties of the model-hardware pair rather than of the agent, which is exactly the shape a carryover story needs.
Distribution is the least speculative part: the public BaseRT repository ships command-line tools, the model format, a public C API and Python, Node, Rust and Swift bindings under Apache 2.0, and it can serve an OpenAI-compatible API [12]. B:OS itself has no disclosed revenue, contracts or pricing, and the experiment does not establish what anyone will pay for automated optimization [10].
The number that would settle this is not a faster kernel. It is a second port whose agents start from the first port's findings, reported as experiments-to-parity. If port two still needs 21 experiments, Base Compute has cheap labor, not accumulated knowledge.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Each new architecture can otherwise require another round of operator implementation, numerical verification and device-specific optimization by systems engineers.
Base Compute gave Claude Fable 5, Kimi K3 and GLM 5.2 the same completed Nemotron port, the same evaluation setup and an eight-hour limit.
Lukas Wesemann, Prabod Rathnayaka and Fabian Waschkowski, the three Base Compute co-founders, published results on August 20 from a system called the Base Optimization Stack (B:OS) that assigns AI agents the low-level work of porting open-weight models to new processors and tuning their kernels.
Base Compute pitches B:OS to model developers, chip vendors and device makers that need a particular model to run efficiently on particular hardware.
In Base Compute's reported results, Claude Fable 5 completed 50 experiments and produced the highest score, for $265.
Kimi K3 ran 21 experiments, reached 97% of Fable's performance, and cost $77.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific numbers, single self-reported run
The disclosure is unusually granular — experiment counts, dollar costs, relative scores, prefill and decode throughput, a perplexity tolerance and unit-test gating — and the protocol is described (identical port, identical evaluation setup, eight-hour cap). But every figure comes from the vendor, which selected the model, designed the protocol and ran the tests, and no independent replication exists. The load-bearing business claim, cross-job carryover, has no measurement at all.
Public runtime, demo-stage usage
There is a real, permissively licensed public artifact — the BaseRT repository with a C API, four language bindings and OpenAI-compatible serving, shipped July 1 — and one first public demonstration on Nemotron. Beyond that, adoption signals are absent by the source's own account: no customers, contracts, revenue or pricing for B:OS, and no third-party usage figures. Scored low on disclosed uptake rather than inferred.
Numbers real, repeatability oversold
The measured claims (throughput gains, agent costs, runtime comparisons) are stated with attribution and specificity, so the gap is not driven by inflated metrics. It comes from the framing around them: a product pitched as making per-chip support economically repeatable, and a mission to 'run AGI on-device', resting on one vendor-run port with zero evidence that findings transfer to the next model-chip pair and no disclosed demand. The reporting itself flags these gaps, which limits the overstatement rather than compounding it.
Vendor-designed benchmark promoting own stack
Base Compute chose the model, designed the evaluation protocol, executed the runs and published results that favour its own runtime against llama.cpp and MLX, while positioning B:OS as a product for chip vendors and device makers. The article's own primary sourcing is company material (a newsroom post and Base Compute's BaseRT release article). That is a strong self-interest structure with no adversarial or third-party check in the cluster.
One publisher, one vendor dataset
Confidence is limited by structure, not detail: a single publisher in the cluster relaying a single vendor's unreplicated results. The internal figures are consistent and the derived arithmetic checks out, so the factual account of what was claimed is reliable; whether the performance and economics hold outside this one Nemotron-on-Apple-silicon setup is not assessable from the supplied material.
leadership
You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.1 distinct publisher
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026