Skip to content

Build1 publisher2 min readPublished

d-Matrix's CTO labels Raptor's 1,000-tokens-per-second figure a simulation

d-Matrix CTO Sudeep Bhoja says Raptor's 1,000 tokens per second per user comes from a simulated 72-card system, with full racks not due until Q4 2027. Whole-system speed, power and cost per request are still unmeasured, so buyers are working from early silicon tests and simulations.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying d-Matrix's CTO labels Raptor's 1,000-tokens-per-second figure a simulation
Photo: letsdatascience.com

What happened

  • Early 3D-DRAM silicon for Raptor measured 0.37 picojoules per bit moved, a figure d-Matrix's Hot Chips presentation classes as I/O energy for data movement.
  • d-Matrix's public slides chart modeled Kimi K3 decode at 988 tokens per second per user with eight users in a batch, falling to 785 with 32.
  • In the configuration Bhoja described, GPUs run prefill on the prompt and d-Matrix accelerators run the decoding steps that generate the response.
  • Raptor is expected to complete tape-out, the handoff of its design for fabrication, by the end of 2026.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Any Raptor cost model built now inherits a simulator's output, and that output moves by about a fifth depending on how many users a buyer assumes share each batch.
  • cost Buyers pay for GPU and d-Matrix capacity together, so the price of each useful response depends on how busy both pools stay once KV-cache transfers are counted.
  • decision A team planning 2027 inference capacity has to choose between waiting for Q4 2027 rack measurements and committing on component tests and simulations.

"Some of our results come from real hardware, some come from modeling, and some are still targets," Bhoja wrote in answers to LDS [5]. He splits d-Matrix's evidence into measurements on early memory silicon, simulations of a future Raptor system, and results from the existing Corsair product [17]. Only the first involves Raptor's own silicon [1]. I think this sorting is good practice for a chip still short of tape-out. A buyer's cost model should keep the three apart.

Raptor stacks memory and compute in one package. The goal is to cut the cost of moving data while supplying the bandwidth fast token generation needs [7]. Divide the 0.37 picojoule figure into the 2 to 3 picojoules per bit Bhoja cites for HBM4 and you get 5.4 to 8.1 times less energy per bit moved [3]. That matches the company's own five-to-eight range [9]. For the ratio to reach the power bill intact, data movement would have to be nearly all of a rack's draw. LDS points out that compute, other memory activity and networking also consume energy [18].

The speed chart models a 72-card system serving Kimi K3 with a one-million-token context [10]. Its sizing assumes 4-bit weights and an 8-bit KV cache [11]. Going from a decode batch of eight to one of 32 cuts per-user speed by 21 percent [1]. Multiplied out, aggregate output climbs from about 7,900 to about 25,100 tokens per second, assuming every user in the batch decodes at the charted rate [2]. A service sold on latency cares about the first number. One sold on throughput cares about the second. For either figure to transfer, the buyer's model has to meet its quality targets at 4-bit weights and run near the charted concurrency. The simulator also has to match a chip that has not yet taped out [6].

The prefill-decode split adds costs that a per-user rate leaves out. Moving the KV cache from GPU to accelerator takes time and energy. Each accelerator's memory caps how many users it can serve at once [15]. The operator also has to size two hardware pools, and idle capacity in either one still costs money [16]. Bhoja's advice is to count only requests that meet speed and quality requirements, and to put idle hardware and data transfers in the bill [4].

The September 10 announcement puts initial availability of the NVIDIA-integrated Raptor rack in Q4 2027 [13]. Bhoja said d-Matrix has not measured end-to-end speed or power on a full Raptor system [3]. His last test for any proposal is which NVIDIA dependencies it keeps [4].

What to watch

  • Completion of Raptor tape-out by the end of 2026, the first point where the simulated speed chart can be checked against fabricated chips.
  • The first end-to-end speed, power or cost-per-request measurements from a full Raptor rack, due no earlier than the Q4 2027 availability date.
  • Whether d-Matrix publishes modeled rates above 32 concurrent users, or figures that account for GPU-to-accelerator KV-cache transfer.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories