Skip to content

Build3 publishers3 min readPublished

Reflection sells Beam to sovereign-AI buyers on a 3-to-4x inference compute claim

Reflection AI says Beam, its 501-billion-parameter open-weight model, needs three to four times less compute than comparable open models. Buyers have only the company's figures until the weights ship later in October.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Reflection sells Beam to sovereign-AI buyers on a 3-to-4x inference compute claim
Photo: semafor.com

What happened

  • Beam is a text-only mixture-of-experts model that activates 23 billion parameters per token and was trained on 23.8 trillion tokens, with a one-million-token context window.
  • Reflection's reinforcement-learning run used 10,500 Nvidia GB300 GPUs for four weeks and produced more than 100 million rollouts.
  • Beam is still in final red-teaming and safety evaluations, and Reflection has opened a waitlist for early access.
  • Reflection is aiming Beam at enterprises and public-sector bodies that want to customize a model and run it on their own infrastructure.
  • RuntimeWire's own benchmark review put Beam near or ahead of some open models on selected coding tests and behind others.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction One report ties the 3-4x figure to compute per generated token and another to compute per problem solved, so a team running long reasoning jobs cannot yet tell which bill to expect.
  • cost Hosting cost follows the full parameter count held in memory, so a buyer sizing on-premises hardware has to budget for the whole model even if the compute saving holds.
  • exposure Reflection's backers, Nvidia among them, stand to gain from demand for both the model and the computing infrastructure that self-hosted deployments of it would need.

The efficiency figure appears in two versions. RuntimeWire, relaying Semafor, reports Laskin saying Beam needs three to four times less computing power than comparable open models to reason through a problem [3]. A dev.to account of the launch attaches the same range to inference compute per generated token [5]. Those are different measurements. Per-token compute in a mixture-of-experts model follows the active parameter count, and Beam uses about one parameter in 22 on each token [22]. A per-problem figure also depends on how many tokens the model spends reasoning before it answers.

Hardware is sized by a different number. The model activates only some experts on each token, but all 501 billion parameters stay loaded in memory [6]. If the weights ship at 16-bit precision, they come to about 1 TB before any context cache [24]. RuntimeWire's review found that Reflection's compute comparison leaves out several costs of serving a model [12]. Neither report states the hardware, batch size or precision behind the three-to-four-times figure.

In my view the ratio transfers to a buyer only if that buyer's serving setup resembles the one Reflection measured, with similar batch sizes and output lengths. Until the technical report defines the comparison [1], I would size hardware for the full parameter count and treat the throughput gain as a company claim.

The training disclosure is more concrete, and the engineering behind it is serious. The four-week reinforcement-learning run works out to about 7.06 million GPU-hours [25]. To grade the rollouts, Reflection ran roughly 1.3 billion sandboxes fed by one million coding, agent-use and science environments [18]. The company calls it one of the largest RL runs reported by an open lab [20]. Those rollouts used contexts of up to 256,000 tokens [18], about a quarter of the advertised window [23].

On capability, Reflection says Beam performs comparably to Z.ai's GLM-5.2 on advanced reasoning benchmarks [4]. By the dev.to account, the company also places it close to Qwen 3.8-Max, concedes that Kimi K3 is ahead on raw capability and argues Beam's edge is inference efficiency [10]. For those scores to predict production results, a buyer's workload has to resemble the benchmark suite.

Laskin told Semafor that customers looking to build sovereign AI systems "don't really have very good options today" [7]. Reflection's wider plan is an "AI factory" for customers that train and deploy their own systems in private clouds, on premises or in restricted environments [13]. Most of the open models leading benchmark tables come from Chinese labs, dev.to notes, and Reflection frames Beam as an open option trained outside that group [19]. Its own comparison set also includes Nemotron 3 Ultra and Inkling, Thinking Machines' mixture-of-experts model [11]. Neither is among the Chinese-lab families dev.to names [19]. Reflection's June financing valued the company at $25 billion pre-money, according to Semafor [15].

What to watch

  • Whether the technical report defines the 3-4x comparison: which models, which hardware, and whether it counts per token or per task.
  • The precision and any quantized variants the weights ship in, which set the memory needed to host Beam.
  • Independent long-context results past 256,000 tokens, the longest context used in the RL run.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories