Skip to content

Invest4 publishers2 min readPublished

Reflection AI pitches Beam as a US open model level with China's GLM-5.2

Reflection AI says its first model, Beam, performs at the level of Z.ai's GLM-5.2 while using three to four times less compute than similar open models. Until outside testers report, those are the company's figures, and they count most with buyers who cannot use Chinese models.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Reflection AI pitches Beam as a US open model level with China's GLM-5.2
Photo: semafor.com

What happened

  • Reflection says Beam is approaching Alibaba's Qwen 3.8-Max, a Chinese model that itself trails the top closed systems from OpenAI and Anthropic.
  • Several large Western companies already use Chinese models to cut the cost of repetitive work such as customer service and routine coding.
  • Reflection, backed by Nvidia, Sequoia and Citigroup, confirmed in June that its latest round closed at a $25 billion pre-money valuation.
  • CEO Misha Laskin said Reflection is working with US and UK government AI evaluators to assess what Beam can do.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • contradiction Semafor reports Beam as being released, while Crypto Briefing says it has not been released or benchmarked publicly, so the GLM-5.2 comparison stands on Reflection's word until outside results appear.
  • exposure Western firms running Chinese models for routine code carry the risk US and UK evaluators found, and a domestic model at claimed parity gives their security reviews a concrete substitute to test.
  • cost Reflection's investors set its price before it had released any model, so Beam's outside test results are the first product evidence behind that valuation.

Laskin's compute figure is the claim a buyer can put a price on. Needing three to four times less computing power per problem [2] means Beam would use between a quarter and a third of the compute a comparable open model spends on the same task [17]. Open models are downloaded and run by the user [7], so a company hosting Beam itself would see that saving in its own compute bill. Laskin's comparison group is "comparable open models" [2]. A buyer would want the figure measured against GLM-5.2 on the same workload, or rather against whichever Chinese model it already runs.

Two buyers are in view, and they apply different tests. The first is the large Western company that moved customer service and routine code onto Chinese models to save money [5]. For that firm, a model that matches GLM-5.2 [3] leaves it where it started, less the cost of switching, unless the compute saving holds against the model it already uses. The second is the government or business building sovereign systems that does not or cannot use Chinese models [8]. "They don't really have very good options today," Laskin said [9]. That buyer picks among closed models and Western open-weight models from Mistral and Thinking Machines Lab [12], a group Reflection says Beam outperforms [15].

I think Beam's commercial case sits with the second buyer, for whom GLM-5.2 is a quality yardstick and the competition is Western. The counter-case is that security findings on Chinese models [6] push the cost-cutters into the second group, in which case Beam wins both at parity. Evidence against my view would be large firms moving customer-service work from Chinese models to Beam before anyone publishes a compute comparison against GLM-5.2.

Outside testing could land in a few places. The US and UK evaluators working with Reflection [13] could confirm parity and the compute saving together. They could find the saving only in coding and agentic tasks, where Laskin says Beam is strongest [16]. Or the comparison could go stale before it is settled: GLM-5.2 came out in June [3], and Qwen 3.8-Max is already ahead of Beam [4].

Reflection calls Beam a "workhorse" [15]. Its own placement puts the model below Qwen 3.8-Max, which trails OpenAI and Anthropic [4]. The larger claim is reserved for the model already in training, which Laskin said would be "much more" powerful and performant [10]. Neither report includes a price for Beam or a plan for charging for open weights.

What to watch

  • Results of the US and UK government assessments Reflection says are under way, and whether they test Beam directly against GLM-5.2.
  • A compute or price comparison between Beam and GLM-5.2 on the customer-service and routine-coding work Western firms moved to Chinese models.
  • Reflection's successor model, which Laskin says is already in training, measured against Alibaba's Qwen 3.8-Max.
Loading claim ledger
Loading source directory links
Loading share composer