Build1 publisher3 min readPublished
Reflection's Beam splits two coding benchmarks with GLM 5.2 before its weights ship
Reflection AI's 501-billion-parameter Beam beats GLM 5.2 on SWE-bench Pro and trails it on Terminal-Bench, by the company's own scores. Its claimed compute saving covers only part of a serving bill, so planning around it as a US-built open model has to wait for outside tests.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Beam is a mixture-of-experts model that activates 23 billion parameters per token, according to Reflection's technical announcement.
- Reflection says a four-week reinforcement-learning run on 10,500 Nvidia GB300 GPUs generated more than 100 million rollouts.
- Kimi K3 and Qwen 3.8-Max also score above Beam on Terminal-Bench in Reflection's own comparison table.
- The Financial Times reports Reflection has raised more than $4 billion since its 2024 launch and is valued at $25 billion.
- Early access is by sign-up, and Reflection has promised Apache 2.0 weights, a technical report and developer tools for later in October.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A team that budgets on the three-to-fourfold saving has to cover the prefill, attention and serving costs itself, and nobody outside Reflection has measured those costs for Beam yet.
- decision Each model wins one of the two coding tests. Choosing between Beam and GLM 5.2 for agent work comes down to which benchmark looks more like a team's own tasks.
- capability If the weights ship under Apache 2.0 as planned, organizations that want a self-hosted US-built model of this size get a permissively licensed option.
Reflection says Beam matches GLM 5.2 on reasoning while using three to four times less inference compute [22]. The estimate draws on Artificial Analysis and DataCurve benchmark data. It excludes prompt prefill, context-dependent attention and serving overhead [23]. That leaves per-token model compute. A sparse model looks good on that measure: 23 of 501 billion parameters is about 4.6 percent of the weights per token [17]. For the ratio to carry over to a real bill, the excluded terms would have to be small for a team's traffic, or shrink by the same factor. Beam is aimed at coding, reasoning and agentic tasks [3]. In long agent sessions, the context-dependent attention term grows with the context.
Memory follows the total count. A server has to keep all 501 billion parameters resident, because the router can send the next token to any expert [4]. Keeping them resident is a serving cost, and serving overhead is one of the excluded terms [23]. RuntimeWire, which reported the scorecard, called the estimate a model-compute comparison and not a measure of what a customer will pay to run the system [23].
The two benchmark gaps point in opposite directions. Beam trails GLM 5.2 by 0.9 points on Terminal-Bench 2.1 [7][18]. It leads by 3.4 points on SWE-bench Pro v1 [8][19]. Both are company-published comparisons, and the company gives no single overall ranking [9]. Reflection printed the rows Beam loses, and I give it credit for that. For a model sold on agentic work, second place on the terminal benchmark by under a point is an awkward spot. Either gap only tells a team something if its own tasks look like the benchmark's.
The training spend behind those scores is large. If every GPU in the reinforcement-learning run worked all 28 days, the run used about 7.06 million GPU-hours [20]. Reflection also says it trained Beam on 23.8 trillion tokens [5]. Semafor reported that the June funding round closed at a $25 billion pre-money valuation, with Nvidia, Sequoia and Citigroup among the investors [13].
According to the Financial Times, open-weight models' share of tokens on Vercel's AI Gateway went from 7 percent in December to 56 percent in August [11]. That is an eightfold rise in share, measured on one gateway's traffic [21]. The FT also reports partnerships with the Pentagon and the US Department of Energy, plus agreements to build models for US allies including South Korea [14]. Laskin co-founded Reflection with Ioannis Antonoglou, another former Google DeepMind researcher [1]. He told the FT that Reflection is seeing demand from organizations looking for sovereign AI systems and unwilling to rely on Chinese models [15].
I'd keep Beam out of capacity plans until outside teams have run the released weights on their own traffic. In my context, which is coding agents with long prompts, the costs the estimate leaves out are the ones I'd need measured first [23]. Waiting costs a few weeks [10].
What to watch
- Whether the technical report released with the weights restates the GLM 5.2 compute comparison with prefill, attention and serving overhead included.
- Independent Terminal-Bench 2.1 and SWE-bench Pro v1 runs on the released weights, to see whether the 0.9-point deficit and 3.4-point lead hold.
- The more powerful follow-on model that Laskin told Semafor Reflection is training.