DeepSeek open-sources Ascend AI chip toolchain with Huawei
TECH

DeepSeek open-sources Ascend AI chip toolchain with Huawei

22+
Signals

Strategic Overview

  • 01.
    On September 30, 2026, DeepSeek open-sourced six software components tailored to Huawei's Ascend AI chips, developed with Huawei's support: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect.
  • 02.
    All six released components correspond one-to-one with tools DeepSeek had previously open-sourced for the Nvidia platform, effectively porting its training stack to a second hardware ecosystem.
  • 03.
    TileLang is positioned as a simpler, high-level programming language alternative to Nvidia's CUDA for writing AI kernels.
  • 04.
    DeepSeek and Huawei jointly built a 128-chip supernode solution based on the Ascend 950 chip, and every TileLang operator currently used in DeepSeek's model training now has a corresponding high-performance implementation on Ascend.
  • 05.
    TileLang-Ascend benchmarks show GEMM workloads at roughly 0.98x hand-written Ascend C performance, vector operators around 0.96x, and cube-vector workloads around 0.95x.
  • 06.
    DeepSeek framed the release around the need for a universal, easy-to-program high-level language to build an independent, self-controlled GPU software ecosystem.

Deep Analysis

Inside the Toolchain: Six Tools Built to Mirror CUDA

The September 30 release is not a single tool but a full stack: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, each ported to run on Huawei's Ascend chips [1]. Notably, all six correspond one-to-one with components DeepSeek had already open-sourced for Nvidia GPUs, meaning the company essentially rebuilt its entire Nvidia-facing training stack for a second hardware target rather than writing something new from scratch [3]. At the center is TileLang, pitched as a simpler, high-level programming language that lets developers write AI kernels without touching Ascend's lower-level instruction set directly [4]. DeepSeek and Huawei paired the software drop with a hardware milestone: a jointly built 128-chip supernode based on the Ascend 950, with every TileLang operator currently used in DeepSeek's own model training now carrying a high-performance Ascend implementation [5]. Early benchmarks suggest the abstraction layer isn't costing much performance - TileLang-Ascend GEMM workloads run at roughly 0.98x hand-written Ascend C code, vector operators around 0.96x, and cube-vector workloads around 0.95x [2]. DeepSeek framed the effort in explicitly ideological terms, stating that building an independent, self-controlled GPU software ecosystem starts with a general-purpose, easy-to-program high-level language capable of pushing hardware to its limits [5].

Why Now: Export Controls and China's Self-Reliance Mandate

This release is the fourth step in a staged rollout, not a sudden move. DeepSeek first open-sourced five Nvidia-facing tools during its February 2025 Open Source Week, then shipped TileLang-Ascend in September 2025, followed by Ascend-targeted DeepSeek V4 kernels in April 2026, before completing the full six-component toolchain now [2][8][9]. The throughline is Beijing's push for AI chips that are 'independent and controllable' as U.S. export controls have steadily reduced both the performance and volume of foreign AI chips available to Chinese firms [7]. A key design goal of the toolchain is lowering the switching cost of that transition: tools like TileKernels offer common APIs and hardware-abstraction layers so developers can move workloads between Nvidia and Ascend with less code rewriting, reducing the engineering tax of adopting domestic chips [1]. In that sense, the release reads less like a one-off open-source drop and more like infrastructure built specifically to absorb the shock of tightening chip controls.

The Skeptic's Case: A Language Isn't an Ecosystem

Not everyone reads this as a CUDA-killer. One analyst quoted in coverage of the release argued that a simpler programming model does not by itself make TileLang an alternative to the much broader CUDA ecosystem, since Nvidia's advantage extends past syntax into years of accumulated libraries, developer tooling, optimization work, and sheer developer familiarity [4]. The published benchmarks - TileLang-Ascend hitting 0.95x to 0.98x of hand-written Ascend C performance - measure individual kernels rather than full system economics, leaving open how the toolchain performs at production scale and cost [2]. Adoption history adds further caution - Chinese firms reportedly bought around 1 million Nvidia H20 chips in 2024 versus an estimated 450,000 Huawei Ascend 910B units shipped, a gap that suggests Nvidia's practical pull with Chinese developers hasn't disappeared just because a software bridge now exists [6].

Second-Order Effects: A Ripple Through China's Chip Demand

The toolchain release appears to be moving actual purchasing behavior. Following DeepSeek's Ascend-focused pushes, ByteDance, Tencent, and Alibaba reportedly approached Huawei about new orders for Ascend 950 chips, and Huawei is said to be planning around 750,000 Ascend 950PR units in 2026, with mass production beginning in April and full-scale shipments in the second half of the year [6]. That demand signal matters more than the code itself in some ways: an open-source toolchain only pays off if there's enough Ascend silicon in the market to run it on, and Huawei's hardware supply has historically been the binding constraint, not DeepSeek's willingness to write software for it.

Historical Context

2025-02
DeepSeek's Open Source Week in February 2025 released five tools built for Nvidia GPUs, the direct predecessors to the components later ported to Ascend.
2025-09
TileLang-Ascend was first open-sourced, giving developers a domain-specific language for high-performance workloads on Huawei processors ahead of the fuller September 2026 toolchain release.
2026-04
DeepSeek released DeepSeek V4 kernels for Ascend ahead of the full six-component toolchain open-source.
2026-09-30
DeepSeek open-sourced the full six-component Ascend toolchain while Huawei detailed its jointly-defined SuperPoD Flex system.

Power Map

Key Players
Subject

DeepSeek open-sources Ascend AI chip toolchain with Huawei

DE

DeepSeek

Chinese AI lab that developed and open-sourced the toolchain (TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, DeepSelect) to port its Nvidia-oriented training infrastructure to Huawei Ascend hardware, reducing CUDA dependence for its own model training.

HU

Huawei

Provided the Ascend AI chip hardware (including the Ascend 950 series) and engineering support for the joint 128-card supernode solution, positioning Ascend as a domestic alternative to Nvidia GPUs for Chinese AI developers.

NV

Nvidia

Incumbent whose CUDA software moat and China GPU market share are directly targeted by this release; its software position extends beyond a programming language into years of accumulated libraries, tooling, and developer familiarity that TileLang alone does not replicate.

BY

ByteDance, Tencent, and Alibaba

Reported to have reached out to Huawei about new orders for Ascend 950 AI chips following DeepSeek's Ascend-related releases, signaling broader Chinese industry demand for domestic chips.

Fact Check

9 cited
  1. [1] DeepSeek Opened the Code. Can Huawei Deliver the Compute?
  2. [2] DeepSeek Expands Huawei Ascend Push With Six Open-Source AI Tools
  3. [3] DeepSeek Ascend Tools Mirror Nvidia-Based Predecessors, Huawei Confirms
  4. [4] DeepSeek Targets Nvidia's Software Moat With Huawei Partnership
  5. [5] DeepSeek Builds for Huawei Ascend
  6. [6] DeepSeek's Huawei-Optimized Model Signals That U.S. Chip Controls May Be Losing Their Bite
  7. [7] China's Drive Toward Self-Reliance in Artificial Intelligence Chips and Large Language Models
  8. [8] TileLang: DeepSeek, Huawei, Ascend, and the Push Beyond CUDA
  9. [9] DeepSeek, Huawei Open-Source Ascend Chip Tools

Source Articles

Top 4

THE SIGNAL.

Analysts

“A simpler programming model does not by itself make TileLang an alternative to the much broader CUDA ecosystem, since Nvidia's software position extends beyond a programming model into years of libraries, developer tooling, optimization, frameworks, and accumulated developer familiarity.”

Unnamed industry analyst
Market commentary via Gurufocus
The Crowd

“DeepSeek and Huawei are building an alternative to Nvidia’s CUDA as China ramps up its domestic AI chip industry. Via Bloomberg: DeepSeek has released an open-source toolkit for Huawei’s Ascend accelerators, including TileLang support optimized for Ascend 950. The tools...”

@@kimmonismus738

“DeepSeek-V3.2 shows: - Chinese chips are rising: Day-0 support for Huawei Ascend & Cambricon; - ML compiler: DeepSeek uses TileLang, letting you write Python → compile to optimized kernels on diverse hardware. E.g., 80 lines of Python can reach 95% of FlashMLA’s (CUDA written...”

@@Yuchenj_UW1325

“DeepSeek has open-sourced its AI infrastructure stack for Huawei Ascend NPUs. The release brings several components to Ascend, corresponding to DeepSeek's existing NVIDIA GPU stack: • TileLang — a high-level DSL for writing optimized AI kernels • DeepGEMM-Ascend — optimized...”

@@Chinazhidx204

“DeepSeek to order 160000 Huawei AI chips over Nvidia”

@u/CleaRSightZ593
Broadcast
Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX)

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX)

Breaking the CUDA Moat: How DeepSeek TileLang is Ending NVIDIA Lock-In

Breaking the CUDA Moat: How DeepSeek TileLang is Ending NVIDIA Lock-In

The Hidden Engine Behind DeepSeek V4 - DeepEP V2 and TileKernels Explained

The Hidden Engine Behind DeepSeek V4 - DeepEP V2 and TileKernels Explained