Science1 publisher3 min readPublished
Positron raises $875M for an inference chip built around commodity LPDDR5X instead of HBM
The Reno startup says its Asimov ASIC carries 288 GB of on-package LPDDR5X and realizes more than 90% of its memory bandwidth. Both figures come from the company, and no independent benchmark accompanies them.
The Scientist · Science desk

What happened
- Positron AI raised $875 million across a combined Series C and Series C-1, valuing the Reno-based inference chip startup at $5 billion post-money.
- Its Asimov ASIC, built on TSMC's N3P process, carries 288 GB of on-package LPDDR5X and can expand to 2.3 TB over CXL.
- Positron says Asimov realizes more than 90% of its memory bandwidth on transformer inference workloads, a figure the company supplies itself.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- contradiction The round is described as a bet on memory capacity, yet the number Positron leads with is 2.76 TB/s of realizable bandwidth.
- exposure Verifying the utilization gap is left to whoever writes the first purchase order, because no third-party measurement is in circulation.
- decision For an operator with an air-cooled hall and a fixed power feed, the accelerator that scores higher matters less than the one the building can actually host.
- constraint A cheaper bill of materials only protects margin if software does not get there first, and an 80% inference cost cut arrived with no hardware change.
The range Positron quotes for its memory advantage, roughly 2x to 29x, splices two comparisons that are not measuring the same thing [7]. At the low end, Asimov's 288 GB of on-package LPDDR5X against the H200's 141 GB of HBM3e works out at 2.04 times [1]. At the high end, 2.3 TB reached over CXL against an H100's 80 GB is 28.75 times [2]. The first figure compares one package with another. The second compares a package against a package plus an expansion bus, and the account of the round reports nothing about what that bus does to latency [21].
The bandwidth numbers carry more of the argument than the capacity numbers do. Positron puts Asimov's realizable bandwidth at 2.76 TB/s, against the H100's 3.35 TB/s theoretical peak, of which Forkast says less than 1 TB/s is typically achieved in practice [10]. That is 82% of NVIDIA's peak figure [3] and more than 2.7 times its working figure [4]. The greater-than-90% utilization claim is Positron's own [9]; the under-30% figure for NVIDIA GPUs on transformer inference is asserted by the publication [8]. Neither arrives with a named model, a batch size, a sequence length or a third-party test [21].
These figures do not give tokens per second on a specific model at a latency an operator would actually sell. Whether cheaper memory becomes cheaper serving depends on that number.
The checkable half of the thesis is physical. Asimov uses an organic substrate instead of CoWoS packaging, air cooling instead of liquid, and standard PCIe Gen6 with CXL expansion [12]. It draws 400W against the H100's 700W, about 43% less per chip [14][5]. HBM, by contrast, needs advanced CoWoS packaging and consumes scarce SK hynix and Samsung capacity [11], while LPDDR5X is available from multiple suppliers [22]. Positron's stated target is brownfield deployment, meaning existing sites that cannot support the thermal and power draw of liquid-cooled GPU racks [13].
The margin case rests on token prices that are already falling. The Silicon Data LLM Token Expenditure Index dropped below $1 per million tokens for the first time in early September [15]. As that number compresses, silicon cost structure decides the margin. That is the case for a cheaper bill of materials. It is also the case for software: the same account credits DeepSeek's V4.1 Flash CED architecture with cutting inference costs 80% with no new hardware at all [17], and OpenAI's Jalapeno chip with showing that labs can design their own silicon [16].
NEA, Atreides Management and Valor Equity Partners are return investors from earlier rounds [20]. Jim Clark, co-founder of Silicon Graphics and Netscape, is leading the Series C-1 tranche [18]. Among the co-leads is SemiAnalysis Capital, the investment arm of Dylan Patel's semiconductor analysis firm [2][19]. At $875M into a $5 billion post-money valuation, the round is 17.5% of the company and puts the pre-money mark near $4.1 billion [6], up from just over $1 billion in February 2026 [3].
What to watch
- An independent tokens-per-second measurement on a named model at a stated latency target would test the greater-than-90% utilization claim.
- Any disclosure of what expanding to 2.3 TB over CXL costs in access latency, which the round's announcement does not address.
- Whether LPDDR5X stays multi-sourced and cheap if inference ASICs begin buying it at GPU-scale volumes.