Skip to content

Build3 publishers2 min readPublished Updated

DeepSeek's Ascend ports swap a CUDA dependency for Huawei's CANN

DeepSeek open-sourced six software components for Huawei's Ascend AI chips on September 30, porting infrastructure it first built for Nvidia hardware. The ports cut the software bill for leaving CUDA, and The Neuron argues Ascend's case now turns on chip supply, reliability and full-scale model performance.

The Engineer · Build desk

Photograph accompanying DeepSeek's Ascend ports swap a CUDA dependency for Huawei's CANN
Photo: yahoo.com

What happened

  • DeepGEMM-Ascend ports DeepSeek's optimized matrix-multiplication kernel library to Ascend and keeps the same API that DeepGEMM uses on other hardware.
  • TileKernels gained Huawei support and now selects either an Nvidia or an Ascend backend behind one set of Python APIs.
  • The work extends TileLang-Ascend, a domain-specific language for Huawei processors that was open-sourced in September 2025.
  • Reuters reported that Huawei supported the work as part of an effort toward what DeepSeek called a more independent computing ecosystem.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Labs with code already written to DeepGEMM or TileKernels can trial Huawei hardware by switching backends before they commit engineers to rewriting kernels.
  • constraint Leaving Nvidia through these libraries makes CANN the toolkit a lab has to trust, debug and track versions of, in place of CUDA.
  • cost Until outside developers measure how much porting the tools save, every lab weighing Ascend has to pay for its own evaluation.
  • decision With less porting work to price, a lab's choice about Ascend depends mostly on whether it can get the chips and run full-size models on them reliably.

Both ports split the stack at the same line. The Python functions a developer calls stay fixed, and the kernel code under them is written separately for each vendor [4][5]. Kernels are the small, optimized programs that do most of the repetitive math in training and inference [13]. That is where hardware-specific work piles up. For code already written against DeepGEMM or TileKernels, moving to Ascend becomes a backend choice and the call sites stay as they are [5].

I think DeepSeek picked the right layer. Nvidia's lead rests partly on roughly two decades of developers programming its GPUs through CUDA [9]. The Neuron describes the switching cost plainly: every incompatible piece of a stack "creates another engineering bill" [17].

The vendor dependency moves down one layer. DeepGEMM-Ascend requires Ascend hardware and CANN, Huawei's Compute Architecture for Neural Networks toolkit [7]. TileKernels lists CUDA for its Nvidia backend and CANN for its Ascend backend [8]. A lab that leaves Nvidia through these libraries trades one vendor toolkit for another. Open source, in this release, means anyone can read the code that calls into Huawei's toolkit [7]. The Neuron's report notes that open source does not automatically mean freedom from vendor dependence [16].

The Neuron wrote that "portability at the API layer is different from parity underneath it" [12]. A matched API tells a lab its calls will run on Ascend. Speed at full model size is a separate measurement, and so is whether a long job stays up. The report argues that kernel benchmarks are only part of the test [14]. I'd expect any published kernel figure to carry over only to labs running similar matrix shapes on similar Ascend hardware and CANN versions.

According to The Neuron, independent developers have not yet measured how much engineering work the tools remove from choosing Ascend [10]. The same report says whether Ascend becomes a dependable alternative depends on chip supply, engineering effort, reliability and full-scale model performance [11]. Supply sits outside the code. The report argues that better software cannot solve a chip shortage [15].

What to watch

  • Independent measurements of how much porting work DeepGEMM-Ascend and TileKernels remove for labs outside DeepSeek.
  • Full-scale training or inference runs on Ascend built on these libraries, with reliability data from long jobs.
  • Evidence on whether Chinese labs adopting the stack can obtain Ascend chips in the volumes large models need.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories