Science1 distinct publisher3 min readPublished
Pathway's preprint reports its post-transformer design nearly matching an entry-level OpenAI reasoner on ARC-AGI-1 while metering far cheaper, which makes this a claim about serving economics rather than capability.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The mechanism sits in where the tokens go. A transformer reads the whole prompt at once, then produces its answer a word at a time, and when it reasons it does that reasoning in language, working through the stages of a problem in strict order [11]. Every stage is billed, because metering is per token and a token is roughly four characters of text, ingested or emitted [13]. Harder problems want longer chains [14], and prompt evaluation grows quadratically, so twice the input costs four times the compute [12]. Pathway calls its alternative a post-transformer architecture [10] and offers the cost figure, not the score, as the result.
Invert the study's ratio and it gets concrete: BDH-CQ meters at about 9% of its comparator, since 1 divided by 11 is 0.091 [1]. The parameter gap is wider than the cost gap. At 150 million parameters [6], against the tens of billions to hundreds of billions typical of frontier systems such as Llama 3 70B and Llama 3.1 405B [7], the model is roughly 467 times smaller than the first and about 2,700 times smaller than the second [2]. Smaller models are cheaper to run and quicker to train [8], which is precisely why the question worth asking is whether the cheapness is architectural or just a side effect of being small.
The score carries footnotes. Almost 30% is a two-attempt figure [3], not a single shot, and plenty of models sit well above it [15]. The comparator's result is characterised only as slightly higher, a "modest accuracy gain" [5], so cost per correct answer cannot be computed from what has been published, and cost per correct answer is what a buyer actually spends. Relative token cost is a ratio inside a metering convention rather than dollars on a serving stack, and as reported there are no latency or throughput figures on named hardware.
ARC-AGI-1 asks a system to induce a rule from a few worked examples of shapes and rotations [4]. That is a clean measurement of one capability, and it is not what drives most production bills, which tend to be long inputs and repeated tool calls. A cheap non-verbal reasoner that has not met those workloads has not been priced against them either.
My read, with its conditions stated: this is an economics claim on a preprint [1], not a capability claim, and the load-bearing part is the untested one, that the same cost profile holds when the architecture is built at larger parameter counts [9]. If it does, Pathway's scientists get the test they want of whether wide adoption would change deployment cost and scale [16]. If the advantage came from smallness, it will thin out as soon as accuracy has to rise.
Ranked by verification strength, evidence, and original report placement.
BDH-CQ scored almost 30% on the ARC-AGI-1 benchmark, successfully solving the equivalent of three out of 10 puzzles in two or fewer attempts.
Scientists at AI company Pathway detailed the technical foundations of the BDH-CQ model in a research paper published Aug. 10 on the preprint server arXiv.
BDH-CQ follows a precursor model known as "Dragon Hatchling," created by the same scientists in 2025 and designed to simulate how neurons in the brain connected and strengthened during the learning experience.
ARC-AGI, a 2019 benchmark, uses nonverbal reasoning puzzles such as rotating a series of shapes to complete a sequence to measure the cognitive ability of AI systems.
Parameters for the most advanced frontier models, such as Meta's open-source Llama 3 70B or Llama 3.1 405B, typically number tens of billions to hundreds of billions.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
product
A satirical scoreboard counts 17 agent escapes that hacked somebody else's company1 distinct publisher
product
Meta owns the models and the data centres, and still pays Microsoft to rent someone else's1 distinct publisher
leadership
Project OT put dates on the layoffs before it put numbers on the agents1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one self-authored paper
Every figure that carries this story — the near-30% score, the 11-to-1 cost ratio, the 150 million parameters — traces to a paper Pathway wrote about Pathway's own model, reaching us through a single publisher. Live Science reports no independent ARC-AGI-1 run, and the comparator it names, GPT 5.6 Luna (Low), is corroborated nowhere in our coverage. The architecture explainer around the numbers is solid and conventional; the numbers themselves have been checked by no one.
Preprint stage, nobody using it
The entire record is a dated arXiv posting and a benchmark pass the authors ran themselves. No customer, no deployment, no endpoint, no published price, no other lab building on the design — and its predecessor, Dragon Hatchling, arrives in the story with no usage history either. This is a research artefact that has not yet met anyone's production traffic.
A token ratio promoted into a running-cost claim
The gap sits between Live Science's headline and Live Science's body. The body says an OpenAI reasoner's slight accuracy edge cost roughly 11 times as much in relative token terms on one puzzle set; the headline says the new model costs up to 11 times less to run, which is a claim about operating a system rather than about one benchmark's metering. The opening paragraph then reaches further, offering nonverbal reasoning as a possible next step toward human-like intelligence, on the strength of a score the same piece admits many models beat. Pathway's own framing about scaling and industry-wide cost savings gets relayed without pushback.
Vendor-run comparison against a named rival
Pathway chose the benchmark, ran both sides of the comparison, named a specific OpenAI model as the loser on cost, and concluded that its own architecture would reshape AI economics if the industry adopted it. That is a paper written to be picked up. On the receiving end, a consumer science publisher has its own reason to lead with artificial general intelligence rather than with token accounting — and it does, before the third paragraph.
Clear on what is claimed, thin on what is true
We can state precisely what Pathway asserts and precisely how little of it has been examined, and that combination is worth something. What holds the number down is structural: one publisher, one self-interested paper, an unverifiable comparator model, a cost metric with no stated definition, and even the preprint's year left implicit in the reporting.