DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia ecosystem

DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia ecosystem

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .

DeepSeek says the tools — for which Huawei provided full support during development — are intended to simplify programming while enabling developers to fully leverage the hardware's performance. The two companies also worked together to optimize computation and communication on a supernode system built around 128 Ascend 950 chips. The work addresses two requirements for running large AI workloads across multiple accelerators: performing calculations efficiently on each chip and moving data quickly enough between them to keep the processors occupied.

The released libraries include DeepGEMM-Ascend, which handles matrix multiplication and other calculations used in DeepSeek's models. It supports BF16, FP8, and FP4 operations and uses the same programming interfaces as DeepSeek's existing DeepGEMM library, allowing developers to retain familiar APIs when moving to Ascend. Meanwhile, DeepEP-Ascend handles communication for training and inference, including routing data to the different experts in mixture-of-experts models and combining their outputs. Both libraries were developed and tested on Ascend 950 hardware.

TileLang provides the higher-level programming layer for writing optimized kernels. DeepSeek described it as offering “a simpler programming model” than Nvidia's CUDA, with the aim of improving development efficiency and simplifying code. The language already supported Nvidia and other hardware, with earlier adapters available for Huawei's Ascend processors. The September 30 update adds native support for Ascend 950, including code generation, automatic scheduling, and synchronization.

The tools build on Huawei's existing CANN software platform , which provides the underlying infrastructure for running AI workloads on Ascend. Nvidia's CUDA platform has long supplied developers with a mature programming environment and libraries optimized for its GPUs, making the software ecosystem a major part of the company's advantage in AI computing. While DeepSeek's release provides developers with additional tools to optimize workloads on Huawei hardware, TileLang's support for multiple platforms ensures the language remains useful for Nvidia GPUs.

China’s Huawei to enter South Korean AI chip market with new Atlas SuperPods, clusters pack 8,192 Ascend 950 accelerators per deployment

Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters

Huawei shelves global AI chip rollout as China's own demand outstrips supply

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment