Product1 publisher3 min readPublished
Positron bets $875m that smartphone memory can carry inference
Positron is now worth $5 billion on a design that swaps HBM for the memory that ships in phones, but Asimov has not taped out, mass production is set for the second half of 2027, and the 26x tokens-per-dollar figure comes from simulation.
The Product Desk · Product desk

What happened
- The appliances use LPDDR5X, the memory variety that mostly goes into smartphones, which is cheaper and easier to obtain than HBM but carries significantly less bandwidth.
- Neither Titan nor the Asimov chip is in production, and the claim of up to 26 times the tokens per dollar of Nvidia's Blackwell GB300 NVL72 comes from Positron's simulations.
- The company expects to tape out Asimov at the end of the year on TSMC's three-nanometer node, with mass production following in the second half of 2027.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint No capacity plan written this year gets out of the HBM queue through Positron, because the part that would do it has not been fabricated and cannot ship in volume before the second half of 2027.
- cost Whoever commits early absorbs the gap between a simulated rack result and measured throughput on first silicon, and that ratio is the entire price argument for buying the appliance at all.
- exposure Positron has taken on a second supply dependency in place of HBM: its ramp rests on LPDDR5X commitments it says it still has to secure from partners.
The utilization claim is the load-bearing part of this design, and it can be sized. Positron says most AI accelerators use less than 30% of their HBM modules' bandwidth, while its own appliances unlock more than 90% of LPDDR5X's throughput [6]. That is a factor of at least three in bandwidth delivered per unit of peak bandwidth bought [2]. The substitution works if LPDDR5X's peak deficit against HBM is smaller than three. Positron published the two utilization percentages and left out the peak-bandwidth comparison against a server card [6].
Titan's specs can at least be checked against each other. Up to 18.4 terabytes of LPDDR5X delivering 23.68 terabits per second [7] is about 2.96 terabytes per second [3], so reading the entire pool once takes roughly 6.2 seconds [4]. No usable token rate survives a full sweep per token, which means the hardware fetches a subset of the stored weights on each pass [6]. That is what the activation-function modules are for, and Positron says they determine which of a model's neural networks participate in an inference task [8].
About ten months separate this month's announcement from the earliest mass production date the company has given [5]. The tape-out has to land first, at the end of the year, on TSMC's three-nanometer node [13]. "Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us," said Chief Executive Officer Mitesh Agrawal [11]. "This financing gives us the resources to do exactly that." [12]
Five times February works back to roughly $1 billion for the company seven months ago [1]. NEA, Atreides Management, Valor Equity Partners, Andra Capital, SemiAnalysis Capital and Silicon Graphics and Netscape co-founder Jim Clark led the round, alongside more than a dozen institutional investors [2]. Positron says its manufacturing ramp will emphasise securing supply commitments from partners for LPDDR5X [13], the part it is designing around and the one that mostly goes into phones [5].
Who this is for turns on where the ceiling sits on a given inference workload: memory capacity, or arithmetic throughput on a model that already fits in the boxes on the floor. It also turns on whether the install date can move to 2028. A buyer in the memory-bound, patient corner is the one Titan is built for, and Positron's pitch there is a single appliance running a 32-trillion-parameter model with a 10-billion-token context window [7]. A buyer whose models fit and whose problem is tokens per second this quarter is still queueing for HBM, and the interconnects that attach it are short too [4].
Positron's 26x is a rack-level simulation result, and the silicon it describes has not been fabricated [10]. Teams under capacity pressure tend to treat a tokens-per-dollar ratio as a procurement fact.
What to watch
- Whether Asimov actually tapes out on TSMC's three-nanometer node by the end of the year, since the 2027 production date is timed off it.
- Measured tokens per dollar on real silicon, to test the simulated 26x against the GB300 NVL72.
- A first named Titan customer, and whether anyone orders at the 16,384-accelerator cluster scale Positron advertises.