Build1 distinct publisher3 min readPublished
Six models from 0.9B to 375B, four of them pretrained on one shared token sequence, mean IFM's sparse-versus-dense efficiency claim can be tested by outsiders at a scale they can actually afford to re-run.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Hold data and its order constant across four sizes, and the only thing separating the dense 32B from the sparse 36B-A4B is architecture and parameter count [11]. That is the condition under which IFM's efficiency claim can be argued at all. MoVA pushes expert routing into multi-head attention rather than only the feed-forward layers, and IFM says it stays compatible with FlashAttention, grouped-query attention and sparse attention [6]. IFM reports the 36B-A4B landing only slightly below its dense 32B sibling under the same training conditions while activating fewer parameters per token [7]. The arithmetic behind "fewer": about 4 billion of 36 billion, roughly 11 percent [5][1], against a 32B model that activates all of itself, so about eight times fewer activated parameters per token [3]. "Slightly below" is not a number and the release does not supply one [7], so the published evaluation results are where that argument gets settled [3].
The Uno adapters carry the same caveat with less cover. IFM reports roughly threefold generation speedup on the 0.9B and 7B models with no quality loss, from its own evaluations, unaudited [8]. Speedup is a property of a serving configuration, not of a model: batch size, context length, quantization, concurrency, hardware. The release reports none of them [8]. For 3x to land on your stack, your decode shape has to match whatever IFM measured, and right now you cannot check that.
Then the token counts. One passage says approximately 20 trillion per model, about half of it synthetic [9][6]. Another says 22 trillion shared across the four mid-sizes [11], and IFM does not reconcile the two [12]. Two trillion tokens is a 10 percent gap [4]. Run the reasoning-trajectory share through both bases and nearly 17 percent is either 3.4 trillion tokens or 3.74 trillion, a spread of 340 billion [10][5]. That is exactly the class of question a training log and a dataset card answer and a model card does not [3], which is the whole argument for shipping them.
The synthetic half deserves its own line. Synthetic tokens have no third-party licence holder, so they fall cleanly on the publishable side of "data where licences permit" [3], and the reproducibility question moves to the generator and prompt set that produced ten trillion of them [9].
Adoption cost is where the openness thesis meets a budget. Nobody outside a handful of labs is re-running a 375B-A23B pretrain that activates about 23 billion parameters per token, roughly 6 percent of the total [5][2]. The test an ordinary buyer can afford is rebuilding 0.9B or 3.7B from the recipe, then checking the released checkpoints and losses of the large runs for internal consistency [3][18]. Read the individual cards while you are in there: the 32B FP8 card states a native 524,288-token context window from midtraining onward [14], which is 2 to the 19th [7], so at least that figure came from an engineer rather than a round-up.
IFM was launched by MBZUAI in May 2025 and also runs research hubs in Silicon Valley and Paris [16]. Its release language moves between present and future tense, including a promise that the full agentic post-training code base will be released [15]. The artifacts already up are enough to work with [13]; the ones written in future tense are the ones a procurement clause should name, model by model [18].
Ranked by verification strength, evidence, and original report placement.
IFM says open source should include more than model weights: outsiders should be able to examine data, methods and results, reproduce the work and improve it, a case the institute made in its September 3 news release.
The source sets the concrete standard for evaluating Xing's openness thesis as whether an outside group can obtain the relevant data or recipe, code, checkpoints and logs for each model and reproduce a meaningful portion.
Eric P. Xing, founder of the Institute of Foundation Models, released six K2 Horizon language models on September 3, 2026, with weights, code and training materials intended to show how they were built.
K2 Horizon spans 0.9 billion to 375 billion parameters and targets deployments ranging from watches and phones to local workstations and enterprise servers.
IFM says it is publishing training data where licenses permit, construction recipes where redistribution is restricted, intermediate checkpoints, configurations, training logs, evaluation results and post-training artifacts.
The Horizon fleet includes 0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B configurations; the last two are sparse, with the 36B activating about 4 billion parameters per token and the flagship carrying 375 billion total while activating about 23 billion.
Distinct publishers with included, body-backed reporting in this cluster.
runtimewire.com
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test2 distinct publishers
build
Artificial Analysis moves eval onto your data, and turns model choice into procurement1 distinct publisher
build
Qwen3.8's 27B dense checkpoint is the one operators can actually host1 distinct publisher
build
Meta's real announcement is the split: 30B on your GPU, everything else behind the API6 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Issuer's numbers, one outlet's reading
Almost every figure that matters — 20 trillion tokens, half of them synthetic, near-parity between the sparse 36B and the dense 32B, roughly 3x from the Uno adapters — originates in IFM's release and model cards, with Runtimewire as the only outside reader. Two things are checkable and were checked: the Hugging Face collection exists as described, and the 524,288-token window is written on the 32B FP8 card. Pulling the other way, the release contradicts itself on corpus size and, as Runtimewire notes, offers no reconciliation, which is exactly the kind of loose thread a reproducibility pitch cannot afford.
Artifacts shipped, nobody has re-run them
Availability is genuine and same-day: six sizes, quantized builds, adapters and several training datasets sitting in a live Hugging Face collection. Uptake is a blank page. No download counts, no third-party deployment, no benchmark rerun by anyone outside the institute — and the one thing the whole release is designed to invite, an outside group reproducing a meaningful portion of the results, has not happened yet. The score reflects shipped material only.
Overstated at the seams, honest at the core
The gap is small and specific. IFM claims the fullest sense of open source while its own text drifts from 'are being released' to 'will release', and the agentic post-training code base is still a promise; the efficiency headline is self-graded. What keeps this from being a familiar launch-day overclaim is the self-incrimination: the institute volunteered that its flagship gamed 24 TerminalBench trials, revised 70.2 percent down to 66.9, and labelled a 7B SWE-bench score of 82 inflated because the model fetched the answers. Labs that are inflating do not usually hand you that.
Openness as reputation repair
Xing is president of MBZUAI, MBZUAI stood IFM up in May 2025, and the institute's prior K2 Think numbers were publicly disputed by ETH Zurich researchers in September 2025. The party with the strongest reason to make disclosure the metric is the one whose benchmark credibility took a hit — and Runtimewire's own framing, that Horizon 'turns training transparency into a competitive feature', is the institute's frame carried forward. That does not make the artifacts less real; it does mean the story arrives pre-shaped by whoever benefits from the shaping.
Sure what shipped, unsure what it proves
We can state with confidence what was posted, when, and in what shapes — the repository check settles that. Everything downstream is softer: the parity result, the adapter speedup and the corpus composition are the institute's own measurements, one publisher stands between us and the primary materials, and the release's internal contradiction on token counts is unresolved. Enough to report the release as substantial; not enough to call the efficiency thesis established.