Invest1 publisher2 min readPublished
Huawei claims 244 gigabytes of shared HBM for each of the Atlas 960E's 4,096 chips
Huawei used HUAWEI CONNECT 2026 to publish its own specification for a 4,096-chip cluster with a petabyte of shared memory, aimed at models approaching 10 trillion parameters.
The Investor · Invest desk

What happened
- Huawei unveiled the Atlas 960E SuperPoD at HUAWEI CONNECT 2026 in Shanghai on September 17, a cluster that fits up to 4,096 Ascend NPUs into a single unified memory-addressing scheme.
- The company claims 8 EFLOPS for FP8 computations and 16 EFLOPS for FP4, with up to 1 petabyte of High Bandwidth Memory across the full configuration.
- Against the previous Atlas 950 SuperPoD, Huawei says the 960E is 2.3 to 4 times better on training and inference for models approaching 10 trillion parameters.
- Huawei is aiming at one AI computing framework scalable to 256,000 nodes and SuperClusters above a million NPUs, with larger Atlas 960 versions expected in the fourth quarter of 2027.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- capability A Chinese lab that cannot buy Nvidia can now size a 10-trillion-parameter run against a published cluster specification. That planning work can start before a pod is installed anywhere.
- cost The million-NPU ambition works out at about 244 of these pods, and 550 kilowatts of saving each is roughly 134 megawatts of grid capacity an operator would not have to contract for.
- constraint Huawei did not disclose a price, a pod's total draw, or when the 4,096-NPU configuration ships, so a buyer can model the training run but cannot yet schedule the purchase.
- exposure A run addressing one memory pool is exposed to the whole pod's downtime, and 0.2% unavailability is about 17.5 hours a year. On a trainer's checklist, the doubled mean time between failures comes ahead of peak throughput.
Divide the claimed petabyte of HBM by the claimed 4,096 chips and each Ascend NPU carries about 244 gigabytes, with roughly 1.95 petaflops of FP8 behind it [1][2]. Run the same division against the target model size and the ratio gets more useful. 10 trillion parameters stored at one byte a weight is 10 terabytes, so the pod would hold about 100 bytes of memory for every parameter in the model Huawei says it is built for [3].
The improvement claim is a range. Huawei puts the 960E at 2.3 to 4 times the Atlas 950 on training and inference at that model scale [4]. 4 divided by 2.3 is 1.74, so a job that fits one pod at the top of the range needs about 1.7 pods at the bottom [11].
The optics substitution is the specification I would want confirmed in an installed rack. Huawei says roughly 48,000 conventional 800G modules come out and about 5,500 Hi-ONE near-packaged optics units go in, which is 8.7 modules retired for every unit added [5][4]. Against the 42,500 modules that leave, a saving of more than 550 kilowatts works out at about 13 watts each [6][5], or about 134 watts per NPU across the pod [6]. At continuous draw, 550 kilowatts is 4.8 gigawatt hours a year [7].
Cryptobriefing.com, which reported the launch, wrote that the cluster was "designed to go toe-to-toe with Nvidia in the AI infrastructure market" [13]. Savings of this size, it argued, "translate directly into total cost of ownership advantages that can shift procurement decisions even when raw performance is close" [11]. The same report says the Ascend 960 development program is running ahead of its original schedule [10]. The Nvidia comparison rests on Huawei's numbers alone.
The specification supports more than one reading. Either the 960E ships close to its published figures, in which case power and reliability settle Chinese procurement before peak throughput does. Or the 4,096-NPU pod is a demonstration, and volume arrives with the larger platform versions Huawei has dated [9]. Or the single address space holds only at configurations smaller than the headline one [2]. On the evidence published so far this is a design target with a conference slot. What would settle it is a named buyer with an installed pod and a completed run at 10 trillion parameters.
What to watch
- A named customer with an installed Atlas 960E pod. That would move this from specification to procurement.
- Whether the larger Atlas 960 versions hold their fourth-quarter 2027 date, given the report that the Ascend 960 program is ahead of schedule.
- A disclosed total draw for a pod, so the 550 kilowatt saving can be read as a percentage of the electricity bill.