Skip to content

Invest8 publishers2 min readPublished Updated

DeepSeek gives away an Ascend 950 version of the toolkit it built for Nvidia chips

DeepSeek open-sourced six modules for programming Huawei's Ascend 950, versions of tools it first built for Nvidia chips. The largest Ascend build tied to the release is DeepSeek's own plan for a data centre holding at least 160,000 of the chips.

The Investor · Invest desk

Photograph accompanying DeepSeek gives away an Ascend 950 version of the toolkit it built for Nvidia chips
Photo: scmp.com

What happened

  • The accompanying libraries, DeepGEMM, FlashMLA, TileKernel, DeepSelect and DeepEP, cover matrix multiplication, memory management and inter-chip communication.
  • The software is tuned for the Ascend 950 and supports a supernode configuration that links 128 of the chips together.
  • The release follows DeepSeek's April 2026 work adapting its V4 model to run on Huawei hardware.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost DeepSeek pays to port and maintain the libraries and charges nothing for them, so any payback has to come through cheaper operation of its own planned Ascend site.
  • contradiction Crypto Briefing's 'no Nvidia required' framing conflicts with a language built to target both vendors, so the switching cost falls for moves back to Nvidia as well as away.
  • decision For Chinese labs already cut off from Nvidia's top chips, choosing Ascend now turns more on chip performance and supply, because kernels no longer start from a blank page.

The money in this release is in the hardware. DeepSeek charges nothing for the tools [6]. The one deployment with a number attached is DeepSeek's own: a planned data centre in Inner Mongolia for at least 160,000 Ascend accelerators [8]. At 128 chips per supernode [5], a site that size would need 1,250 supernodes if every chip sat in one [1].

I think DeepSeek gets the first benefit. Porting its own Nvidia-era libraries to Ascend [2] makes that fleet cheaper to run, and releasing them free [6] lets outside developers test and fix them. The work extends a port from April 2026, when DeepSeek adapted its V4 model to run on Huawei hardware [9]. Huawei reportedly gave extensive support during development [7]. According to its WeChat post, DeepSeek wants an "independent and controllable" software ecosystem [11].

DeepSeek has also kept Nvidia in the stack. According to the project's GitHub page, TileLang lists Nvidia as its primary back end and now officially supports the Ascend 950 with "native code generation, automatic scheduling, and synchronisation" [3]. Crypto Briefing described the release as "a full software stack purpose-built for Huawei's chips, no Nvidia required" [13]. The GitHub listing describes something narrower: one language that generates kernels for either vendor's chips. That makes it cheaper to move kernel code off CUDA, and just as cheap to move it back.

For Nvidia in China, a portable kernel layer puts more weight on chip performance and supply. US export controls already restrict Chinese access to Nvidia's most advanced parts [12].

Chinese labs could start writing new kernels in TileLang. Ascend would then become a hardware choice for them, and Nvidia's software edge in China would shrink to the CUDA code already written. Use could instead stay mostly inside DeepSeek and Huawei. In that case the release is a running-cost saving on one planned site, published in the open. A third possibility is that the libraries run but the Ascend 950 trails Nvidia chip for chip, and the benchmarking tools shipped with the release [10] make that gap easier to measure. Neither report includes performance figures or a count of outside users.

For now I'd put the most weight on the second outcome, because DeepSeek's own site is the only deployment either report names [8]. Crypto Briefing called the free release a deliberate choice to lower the barrier for developers who might otherwise default to Nvidia because switching felt too costly [14]. That case holds if other Chinese labs ship models trained on Ascend with these libraries. It fails if the Ascend kernels keep coming only from DeepSeek and Huawei engineers.

What to watch

  • Whether labs other than DeepSeek publish Ascend kernels or models built with TileLang and DeepEP.
  • Benchmark results from the bundled tools comparing Ascend 950 supernodes with Nvidia parts on the same kernels.
  • Construction, financing and timing details for DeepSeek's Inner Mongolia site of at least 160,000 Ascend chips.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories