Skip to content

Topic

Open-Weight Model Publishing

Who publishes open model weights, at what scale, and with what portfolio strategy across size bands.

Current stories

build5 publishers

Aleph Alpha's open Kolibri model routes each token through 3.46B of its 78.1B parameters

Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.

Perspective Coverage

5 publishers
Builder
Builder 49%
Operator
Operator 36%
Investor
Investor 15%

Reality

Evidence70
Adoption
Insufficient
Hype gap+15
Incentives65
Confidence68
build2 publishers

Reflection AI commits $150 million a month to SpaceX compute before shipping its first open model

Reflection AI is preparing its first open-weight model after signing compute deals worth $150 million a month with SpaceX and over $1 billion with Nebius. Open weights leave the infrastructure and upkeep with the customer, so how the model deploys is the test buyers can run themselves.

Reality

Evidence45
Adoption8
Hype gap+25
Incentives60
Confidence50
invest1 publisher

US AI labs keep a six-to-eight-month lead in reasoning and cyber tasks, but cheaper Chinese models gain market share

Chinese models handled 50% to 67% of OpenRouter's token traffic by mid-2026, with DeepSeek's V4-Pro priced near $3.96 per million output tokens. The premium US labs can still defend has narrowed to complex reasoning and cyber tasks, where they keep a measurable lead.

Reality

Evidence35
Adoption50
Hype gap+25
Incentives
Insufficient
Confidence35
build1 publisher

Moonshot publishes open weights for its trillion-parameter K2.7-Code model

Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence40
product3 publishers

AWS open-sources Strands Decider 2B to make agent routing decisions on a laptop

AWS open-sourced Strands Decider 2B, a model that picks from a fixed list of options in under 150 milliseconds on a local machine, the company says. Agent builders get a small model they can run themselves for routing and tool-selection steps that would otherwise each call a full LLM.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 25%
Investor
Investor 27%

Reality

Evidence58
Adoption
Insufficient
Hype gap+22
Incentives60
Confidence62
build3 publishers

Typesafe's Jev API gains an open-weight rival in Cloudflare's Clef decision models

Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.

Perspective Coverage

3 publishers
Builder
Builder 53%
Operator
Operator 27%
Investor
Investor 20%

Reality

Evidence50
Adoption18
Hype gap+35
Incentives70
Confidence60
invest8 publishers

OpenAI says it disrupted Moonshot-linked users trying to extract its models' protected reasoning

OpenAI says it disrupted a large-scale effort by users tied to Moonshot AI to extract protected reasoning from its models. That makes three US accusers of the lab behind Kimi K3, and the published evidence so far covers attempted extraction only.

Perspective Coverage

8 publishers
Builder
Builder 33%
Operator
Operator 35%
Investor
Investor 32%

Reality

Evidence55
Adoption
Insufficient
Hype gap+35
Incentives65
Confidence60
product1 publisher

Anthropic's weight edit drops GLM-5.3's refusal scores from about 90% to as low as 2%

Anthropic researchers edited the weights of Z.ai's open-weight GLM-5.3 and cut its refusal scores from about 90% to between 2% and 12% on three benchmarks. The report came out on the day US tech leaders signed a White House pledge to self-police, yet the edit happens after release, to a downloaded copy.

Publishers:gizmodo.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence40
build3 publishers

Kimi K3 on its cheapest host undercuts Fireworks' Ember-1 despite a 23% cut in reasoning tokens

Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 30%
Investor
Investor 18%

Reality

Evidence55
Adoption30
Hype gap+25
Incentives70
Confidence58
invest13 publishers

Mistral's record round covers three quarters of the data centres it already ordered

Samsung led the 3bn euros at a mark above 21bn, ASML extended High NA commitments to Samsung and TSMC on schedules running out to 2033, and neither disclosure says how large Samsung's cheque actually was, which is the number this turns on.

Perspective Coverage

13 publishers
Builder
Builder 24%
Operator
Operator 25%
Investor
Investor 51%

Reality

Evidence62
Adoption40
Hype gap+25
Incentives72
Confidence60
invest27 publishers

Nvidia's $12.93 billion Hugging Face deal works out to about $4,300 per shared model

The $12.93 billion price works out to roughly $4,300 per model and $862,000 per integrating enterprise. The asset those numbers describe is the default place teams go for weights they could copy anywhere.

Perspective Coverage

27 publishers
Builder
Builder 28%
Operator
Operator 29%
Investor
Investor 43%

Reality

Evidence72
Adoption65
Hype gap+15
Incentives70
Confidence70
leadership3 publishers

Guardrails that blocked Hugging Face's responders put AI access on the incident plan

Commercial AI models refused every query from Hugging Face's breach responders, Veracode's Chris Wysopal wrote, forcing them onto a self-hosted Chinese model. Security leaders now have to settle which AI model their responders can use before an intrusion starts.

Perspective Coverage

3 publishers
Builder
Builder 22%
Operator
Operator 45%
Investor
Investor 33%

Reality

Evidence45
Adoption
Insufficient
Hype gap+30
Incentives55
Confidence50

Earlier coverage

  1. Recorded Future singles out deepfakes as the AI phishing tactic existing controls fail to stop

    Security · September 29, 2026 · 1 publisher

  2. SuperWhisper's open-weight S1-mini cleans up raw speech transcripts offline on a laptop CPU

    Science · September 29, 2026 · 1 publisher

  3. Mistral takes 3 billion euros from a Samsung-led round to build compute it owns in Europe

    Build · September 27, 2026 · 1 publisher

  4. Supersonic Labs trained Julia 1 for $104 on cloud GPUs, then benchmarked inference on a laptop CPU

    Build · September 27, 2026 · 2 publishers

  5. Washington drafts a letter turning 35 AI signatures into a forced choice

    Product · August 15, 2026 · 1 publisher

  6. Kimi K3's real gate is 1.4TB of VRAM and a bespoke licence, not engineering

    Build · August 20, 2026 · 2 publishers

  7. The open-weight default is for sale at $13B, and its buyer pool ships models too

    Product · August 23, 2026 · 4 publishers

  8. Harvey builds its flagship legal model on Chinese open weights, not its investor's API

    Product · August 23, 2026 · 2 publishers

  9. Apple's $18,299 Mac Studio versus a $200 subscription: 91 months to break even

    Leadership · August 25, 2026 · 2 publishers

  10. Three deals in weeks pull the open-weight distribution layer inside vendor stacks

    Product · August 28, 2026 · 8 publishers

  11. Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.6

    Product · September 3, 2026 · 3 publishers

  12. Spark 1.3's index jump lands on the three tests that carry half the score

    Build · September 3, 2026 · 6 publishers

  13. Nvidia's $12.9 billion buys the model registry 200,000 companies deploy from

    Build · September 3, 2026 · 9 publishers

  14. Chinese banks and telcos are retailing AI tokens in a unit their customers cannot price

    Product · September 5, 2026 · 2 publishers

  15. NVIDIA's Time of Ownership clock runs while the early part waits for the late one

    Build · September 10, 2026 · 3 publishers

  16. Harvey's $550M round prices legal AI at 39 times the ARR it just crossed

    Invest · September 9, 2026 · 2 publishers

  17. Garry Tan points regulators away from distillation days after agencies named six Chinese labs

    Leadership · September 11, 2026 · 2 publishers

  18. DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

    Build · September 11, 2026 · 3 publishers

  19. Salesforce's Koa model reaches pilots with two accounts of what it was trained on

    Product · September 15, 2026 · 3 publishers

  20. Ninety-five dollars of rented H100 time produced Kev's three-model Qwen3.5 port

    Build · September 21, 2026 · 2 publishers

  21. Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

    Invest · September 22, 2026 · 16 publishers

  22. ElevenLabs' CEO will squeeze margins for share in a voice market he expects to even out

    Product · September 24, 2026 · 1 publisher

  23. Hugging Face turned to a Chinese open-weight model to investigate an OpenAI model's breakout

    Science · September 24, 2026 · 1 publisher

  24. Anthropic's open-weights holdout turns a safety argument into a supply question

    Leadership · September 22, 2026 · 2 publishers

  25. NVIDIA's 3D CT model leaves clinical validation to whoever post-trains it

    Build · September 23, 2026 · 1 publisher

  26. Choosing US-only inference for Kimi K3 costs about 11% more than Bedrock's global profile

    Build · September 23, 2026 · 1 publisher

  27. DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

    Build · September 22, 2026 · 1 publisher

  28. Kaitchup traces Bonsai 2's 98.2% retention figure to unpacked weights on an H100

    Build · September 22, 2026 · 1 publisher

  29. China builds AI capacity one phase a year in Inner Mongolia's potato capital

    Product · September 22, 2026 · 1 publisher

  30. Xiaomi's MiMo-V2.6-Pro leads the open-weight index at $0.87 per million output tokens

    Product · September 22, 2026 · 1 publisher

  31. VoxCPM2 trades the per-character speech bill for a GPU and a base-URL change

    Build · September 21, 2026 · 1 publisher

  32. Heretic automatically strips refusals from many open-weight language models

    Security · September 21, 2026 · 1 publisher

  33. Harvey's cost of serving a dollar of revenue tripled in six months

    Invest · September 21, 2026 · 1 publisher

  34. Nvidia spends $7bn to put a free model on the chips it already sold

    Product · September 21, 2026 · 1 publisher

  35. Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price

    Build · September 20, 2026 · 2 publishers

  36. Hacktoberfest's spam fix moves the unpaid hours from PR reviewers to Fest organizers

    Build · September 20, 2026 · 1 publisher

  37. ZCode encrypted a developer's Git history with a key only Z.ai's servers hold

    Product · September 20, 2026 · 1 publisher

  38. Mozilla put Mistral's models behind Firefox Smart Window under zero data retention

    Build · September 20, 2026 · 1 publisher

  39. Open-weight models took 78.4% of Vercel gateway tokens on a single September day

    Build · September 20, 2026 · 1 publisher

  40. Bessent brings AI model rules and rare-earth talks to New York meeting with China

    Build · September 20, 2026 · 1 publisher