Skip to content

Build1 publisher2 min readPublished

d-Matrix models 100 TB/s per card for Raptor, a stacked-DRAM chip due in late 2027

d-Matrix's ISCA paper models up to 100 TB/s per card for Raptor, a stacked-DRAM inference accelerator. The figures come from modeling and the first cards are expected in Nvidia MGX racks in late 2027, so they have no bearing on inference hardware bought in the next year.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying d-Matrix models 100 TB/s per card for Raptor, a stacked-DRAM chip due in late 2027
Photo: letsdatascience.com

What happened

  • The reported Raptor configuration holds 32GB of memory, far less capacity than some HBM-based accelerators.
  • d-Matrix says Raptor bonds a TSMC 4-nanometer logic die face-to-face with a custom DRAM die.
  • In September, d-Matrix said it would integrate its next-generation accelerators into Nvidia's MGX rack architecture using NVLink Fusion.
  • d-Matrix expects Raptor to tape out before the end of 2026.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Per-card bandwidth overstates what a single Raptor card can serve, because large models have to be spread across several cards or into other memory in the rack.
  • constraint Buyers would get Raptor as a decode tier beside Nvidia GPUs doing prefill, so integration work, software and rack economics are judged along with the chip.
  • decision Teams sizing 2027 decode capacity can track Raptor now, but a price-performance comparison has to wait until production cards are measured against shipping accelerators.

The strongest material in the Raptor paper is the engineering around the stack. Autoregressive decoding moves model weights and key-value cache data again and again, and that traffic is what the design targets [5]. I would want power, heat and error rates answered before trusting a bandwidth figure from any part that bonds DRAM to logic. The paper covers power, thermal management and error correction, alongside stream-blocking, which maps KV-cache streams across configurable DRAM channels [6].

The figures resurfaced in an October 2 post on X [3], and reposting a model does not update it. Both throughput baselines are modeled HBM and SRAM configurations. So the 4.71x result does not show a Raptor card beating any specific Nvidia or AMD product in a customer deployment [8]. Dividing the two published ratios puts the SRAM configuration at about 1.93 times the HBM one inside the same model, so the 4.71x headline is measured against the weaker baseline [1]. A multiple carries over to another operator's racks only under conditions. The model mix has to resemble the paper's list: Llama 3.1 70B, DeepSeek-V3, Kimi K2, GPT-OSS, Whisper and Canary [7]. The serving setup has to match the paper's modeled deployment assumptions [2]. Production silicon has to behave like the model.

Capacity is the weak spot. At 8 bits per weight, Llama 3.1 70B's weights come to about 70GB, roughly 2.2 times the 32GB on one card [2]. That means at least three cards for the weights before any KV cache is stored [2]. It is still unclear how work divides between Raptor's local memory and other memory in a rack [14].

I think decode is the right job for a part with high bandwidth and modest capacity. The September arrangement with Nvidia would let an MGX rack assign compute-heavy prefill to GPUs and latency-sensitive decoding to d-Matrix accelerators [11].

d-Matrix said in November 2025 that it had raised $275 million in a Series C round valuing the company at $2 billion, lifting its total disclosed funding to $450 million [15]. Initial availability in MGX racks is expected in the fourth quarter of 2027 [13]. That quarter opens 15 months after the ISCA run in Raleigh ended [3].

What to watch

  • Measured results from production Raptor cards against a shipping Nvidia or AMD accelerator, replacing the modeled HBM and SRAM baselines.
  • A per-card price and yield or volume figures from d-Matrix after tapeout.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories