Skip to content

Build1 publisher3 min readPublished

Ai2's Olmo-core 3 gets 2.7x its old MoE throughput by keeping experts resident on GPUs

Ai2 released Olmo-core 3, an open mixture-of-experts training stack it has benchmarked at over one trillion total parameters. Its speedups are measured against Ai2's own earlier FSDP code, so teams on Megatron-Core must run their own comparison.

The Engineer · Build desk

Illustration accompanying Ai2's Olmo-core 3 gets 2.7x its old MoE throughput by keeping experts resident on GPUs

What happened

  • Olmo-core 3 replaces Ai2's FSDP-based MoE code, which gathered and resharded weights for every small batch, with a DDP design that keeps experts on GPUs and routes data to them.
  • In a preliminary test on eight NVIDIA B300 GPUs, a 47B-parameter MoE ran at 52,000 tokens per second per GPU on the new stack against 19,400 on Ai2's old one.
  • The stack supports the MXFP8 low-precision format, which Ai2 says helps only when compute and data-movement savings exceed the cost of converting between formats.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams already training on Megatron-Core have to benchmark Olmo-core 3 on their own models before migrating, because Ai2 measured its speedup only against its own earlier code.
  • constraint Smaller labs still need enough GPU memory across a cluster to hold every parameter and its training state, so access to hardware decides who can take the trillion-parameter path.
  • capability Academic teams can run expert-count experiments at near-fixed per-token compute on open code without first writing their own expert, pipeline and optimizer sharding.
  • precedent Olmo-core 3 sits behind Ai2's next Olmo generation, so outside teams get the same training infrastructure Ai2 uses for its own models.

Gathering weights for every batch is a poor fit for sparse models. In the 128-expert configuration Ai2 benchmarked, each token is routed to four experts [4]. That is about 3% of the pool [1]. An FSDP setup configured to gather and reshard weights for every small batch [5] therefore assembles expert weights that most tokens never use. Olmo-core 3 moves the data instead. Each GPU permanently holds part of the expert pool, and routed tokens travel to it [c5, c8].

The 2.68x ratio on eight B300s [2] needs its baseline next to it. Ai2 compared the new stack with its own earlier FSDP implementation and called the test preliminary [6]. A 2.7x gain over code that regathered weights every batch says as much about the old configuration as the new one. The post names NVIDIA's Megatron-Core as an established option for training large MoEs [7], but it does not include a throughput comparison against it [c6, c7]. For the number to carry over to another team, that team would need to be running something close to Ai2's old setup, at a model size near 47B, on similar hardware.

The expert-scaling run is stronger evidence for the design. With active parameters held near 3.2B, the share of the model each token uses fell from about 70% at 8 experts [3] to under 7% at 128 [4]. Total size grew roughly tenfold [5]. For throughput to hold up across that range, routing overhead has to stay small as the pool widens. Ai2 itself names that overhead, the communication and coordination cost of sending inputs to experts across a cluster, as what erodes the MoE advantage as models grow [11].

The fixes aimed at that overhead are careful work, and they are the parts I would most want to read in the code. GPU-resident routing keeps routing metadata on the GPU. The CPU can then queue the next work without waiting for that metadata to be copied back [9]. Rowwise expert parallelism writes routed data straight into expert input buffers, so less work goes into rearranging it [9]. Grouped GEMM combines many small expert computations so the GPU can execute them more efficiently [9].

Ai2 pitches the release partly on access. It writes that compute costs put advanced model development out of reach for many academic researchers and smaller labs [12]. The open code saves such a team from building its own routing and parallelism layer [c8, c9]. The memory requirement stays with the team. Ai2 notes that the full MoE still has to be stored across GPU memory and updated during training [11]. Expert parallelism, pipeline parallelism and a distributed optimizer split the model and its optimizer state across GPUs [8]. A trillion-parameter run therefore still needs enough GPUs to hold a trillion parameters plus their training state.

In my view the decision depends on what a team runs today. A group training MoEs on an FSDP setup that regathers weights per batch has a measured reason to try Olmo-core 3 [c5, c6]. A group on Megatron-Core has to run the comparison itself [7].

What to watch

  • A head-to-head throughput comparison between Olmo-core 3 and Megatron-Core on the same model and GPUs.
  • Throughput and GPU-count figures for the run Ai2 benchmarked at over one trillion total parameters.
  • Ai2's MXFP8 measurements, which show whether the format's savings beat its conversion cost on these models.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories