Product1 distinct publisher3 min readPublished
The startup says its weight-space link between GLM-5.2 and Qwen-3.5 costs five percent of the big model while landing midway on quality, which is either a bargain or a downgrade depending on where your acceptance bar sits.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The person who has to act on this owns a routing table: the config that decides which requests go to the expensive endpoint and which get answered locally. For that person the interesting figure is not the midpoint. It is what sits at each end of the line the midpoint divides.
Start with the money, because it is the only unambiguous part. One-twentieth of the cost is a 95 percent saving [12], which means the budget that buys one call to the full GLM-5.2 buys twenty bridged ones [14]. The parameter gap being spanned is roughly 188 to 1 [13]. Those are good numbers on their own terms.
The quality claim is softer than it looks. Halfway is a position on a line, and Wired's account does not say which benchmark or task drew the line [15]. Between a 753-billion-parameter model and a 4-billion one, the distance could be twenty points on something you care about or it could be two. Halfway across a narrow gap is close to free money. Halfway across a wide one may still land below the level you are allowed to ship.
Karl Tuyls, formerly of Google DeepMind and familiar with the work, describes the result as approaching large-model quality without the large model handling the entire loop, with a smaller model running alongside [8]. That is worth reading literally. The frontier model is still in the loop, it is just doing less of it. The pitch is models talking without words, and in practice that means shedding work off the expensive model onto a cheap one without paying for a text round trip in between [2].
There is also a boundary condition sitting inside the mechanism. The bridge operates on values found in the weights [1], the demonstration used two Chinese open-weight models [3], and Wired frames the payoff as making open-weight models more competitive against the closed offerings of labs like Anthropic and OpenAI [11]. Read together, that points away from the endpoint most teams are actually billed for. Most teams hold an API key, not the weights the bridge needs.
The evidence is thinner than the enthusiasm around it. The ARC-AGI 3 result arrives with no details, because Mostik wants to win the contest [5]. Vladimir Arustamian of Lovable, who knows the team, says they have something running that he would have guessed was years away [9], while Mostik's own chief scientist, the Fields medalist Stanislav Smirnov, says there is still no appropriate mathematical language for finding common ground between two models [10]. None of the people quoted here have reproduced the numbers themselves.
The forcing function is a two-by-two you can build from your own logs. One axis is the width of the gap between your small model and your frontier model on your evaluation set. The other is whether your acceptance threshold falls above or below the midpoint of that gap. Wide gap with a threshold under the midpoint is where a bridge pays, and it pays well. Narrow gap means you were overpaying for the frontier model already and should simply drop down. A threshold above the midpoint means the discount buys you answers you will reject, however wide the gap. A gap you have never measured means you are pricing a midpoint between two numbers you do not have.
So run both endpoints against your own set, then price the middle. Sasha Malysheva's bet is that capability will come from combining models rather than scaling a monolithic one [7], and if she is right this becomes ordinary plumbing within a year. The cost of going early is that you inherit a quality level defined by somebody else's benchmark.
Ranked by verification strength, evidence, and original report placement.
Mostik is a Russian startup whose name is the Russian word for bridge; its approach lets different models interact using the mathematical values found in their weights, the values that determine how a prompt becomes an output.
Mostik's method lets AI models talk to one another without producing text output; typical ensembling instead feeds the output of one model into another, which takes a good chunk of time and money.
To demonstrate the idea, Mostik built a bridge between two Chinese open-weight models: the largest version of GLM-5.2, at 753 billion parameters, and a 4-billion-parameter version of Qwen-3.5 that can run on a mobile device.
Mostik CEO Sasha Malysheva, who developed the approach, said it is well known in machine learning that ensembles of models perform better than individual ones.
Malysheva said she personally does not think there will be a monolithic model in future, or that the capabilities of models will come from scaling.
Karl Tuyls, a former Google DeepMind computer scientist familiar with Mostik's technology, said the technique means you can approach large-model quality without the large model handling the entire loop, giving substantial improvements with just a smaller model running alongside, and that it is a no-brainer for anyone tasked with running models as efficiently as possible.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Agents that left notes for each other: inside the 17,600-incident Hugging Face intrusion1 distinct publisher
leadership
A Government Switched Off Two Frontier Models. Your Board Will Want The Fallback Plan.1 distinct publisher
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One coffee meeting, no artifacts
Every load in this story is carried by a single conversation Will Knight had with the founders. Wired relays the twentieth-of-the-cost figure and the exact-midpoint quality figure as Mostik states them; there is no paper, no repository, no eval name, and the one result that could be checked from outside — the ARC-AGI 3 entry — is being withheld on purpose. The two corroborating voices are informed rather than independent: one is a former DeepMind researcher familiar with the tech, the other knows the team.
Nothing yet to count
There is no third party running this. The GLM-Qwen pairing is a demonstration the company assembled to explain itself, and the leaderboard entry is anonymous by choice — so no customers, no downloads, no pricing, no deployments appear anywhere in the reporting. That is not weak adoption, it is adoption that has not been measured.
Framing outruns what a reader can check
"Machine telepathy" and a model that "rocketed to the top" of a notoriously hard competition are doing heavy lifting for a system whose quality midpoint has no benchmark behind it and whose best result is under embargo by the company's own choice. The underlying idea — routing between models through their weights instead of through generated text — is genuinely interesting, and the dek is honest that halfway can read as bargain or downgrade. The distance is between the vocabulary and the verifiable, not between the idea and plausibility.
Everyone quoted has a stake or a tie
Follow who gains from each sentence. Mostik supplies all the numbers and is sitting on its strongest result to protect prize money. Vladimir Arustamian is described as knowing the team. Karl Tuyls is "familiar with the company's tech." Wired discloses each of those relationships, which is the correct move — but after the disclosures, no disinterested party is left in the story, and the founder profile material sits comfortably alongside the technical claims.
Thin record, not doubted reporting
One publisher, one reporter, one sitting. A single independent reproduction or one named eval would move this read a long way in either direction, and right now there is nothing to triangulate against — so the low number reflects how little of the record exists, not scepticism about Wired's account of what it was shown.