DeepSeek-Huawei Ascend Chip Software Partnership Challenges Nvidia CUDA
TECH

DeepSeek-Huawei Ascend Chip Software Partnership Challenges Nvidia CUDA

22+
Signals

Strategic Overview

  • 01.
    DeepSeek open-sourced a six-component software stack for Huawei's Ascend AI chips - TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect - covering matrix math, distributed communication, attention processing, and top-K selection.
  • 02.
    DeepSeek and Huawei jointly built a 'supernode' system linking 128 Ascend 950 chips, tuned to balance computation and data movement across the cluster.
  • 03.
    DeepSeek's V4 model (1.6 trillion parameters) received 'day zero' support on Huawei Ascend 950PR/950DT chips at launch, alongside day-zero adaptation from Cambricon and Hygon.
  • 04.
    DeepSeek is reportedly planning to deploy more than 160,000 Huawei Ascend accelerators at a data center in Inner Mongolia, reducing its reliance on Nvidia hardware for serving its models.

Deep Analysis

Six Tools, One Goal: Replacing CUDA's Developer Experience on Ascend

The release is not one tool but six: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, covering matrix math, distributed communication, attention processing, and top-K selection for Huawei's Ascend NPUs [1]. The centerpiece is TileLang, pitched explicitly as a simpler way to write 'kernels' - the small, performance-critical programs that run directly on chip hardware - without requiring the deep hardware expertise CUDA historically demanded [2]. DeepSeek's own framing is pointed: 'To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential.' On the performance side, industry coverage citing GitHub reporting claims FlashMLA reached roughly 95 percent of theoretical performance on Ascend 950 for some processes [3]- a notable efficiency claim, though one that measures software optimization rather than a head-to-head comparison against Nvidia's newest silicon. The toolkit also underpins a jointly engineered 128-chip Ascend 950 'supernode,' tuned by DeepSeek and Huawei together to balance computation and data movement across the cluster [4].

Why Open-Sourcing Beats Hoarding: Attacking CUDA's Developer Lock-In

The strategically interesting choice here is not that DeepSeek optimized for Ascend - it's that the company gave the tooling away for free. Nvidia's CUDA platform is estimated to have roughly four million developers worldwide, a lock-in built on habit, debugging tools, documentation, and years of accumulated code, not just raw hardware performance [2]. A closed, proprietary toolkit would only ever serve DeepSeek's own workloads; an open-source one invites every developer frustrated by chip export restrictions to start building muscle memory on Ascend instead of CUDA. That is close to the same playbook Nvidia used to make CUDA indispensable in the first place, now aimed back at it. On YouTube, one widely-watched technical breakdown of the release explicitly framed it as a 'two-front attack' on Nvidia - pairing Huawei's hardware with DeepSeek's software credibility - while cautioning that a new stack only matters once developers actually adopt it, not the day it ships. Reaction on X skewed more straightforwardly enthusiastic about the CUDA challenge itself, without that same adoption caveat.

Export Controls, Reframed: The Geopolitics Behind the Code

The backdrop is the US restriction on advanced Nvidia chip sales to China. DeepSeek's founder has said that chip shipment bans, rather than financing, are the company's core constraint, which helps explain the pivot toward Huawei's domestic hardware and software stack [5]. CSIS's analysis frames the stakes sharply: if DeepSeek applied its CUDA-bypassing optimization techniques to strengthening Huawei's Ascend chips and CANN software ecosystem, 'it would pose a much more significant threat to Nvidia' than DeepSeek's model releases alone ever did [5]. That warning has already reached Nvidia's own leadership - CEO Jensen Huang said publicly that DeepSeek optimizing for Huawei's Ascend chips instead of American hardware would be 'a horrible outcome for the United States,' arguing that if AI diffuses globally on a Chinese tech stack, China could become superior to the US in the field [6]. Earlier CSIS research had already documented Huawei's prior-generation Ascend 910C delivering roughly 60 percent of Nvidia H100 inference performance despite export controls - the baseline this new 950-series collaboration is trying to leap beyond [5].

Scale Check: From Day-Zero V4 Support to a 160,000-Chip Buildout

This didn't start with the September toolkit release. In April, DeepSeek's V4 model - a 1.6-trillion-parameter mixture-of-experts model - launched with 'day zero' adaptation support on Huawei's Ascend 950PR and 950DT chips, announced via livestream hours after the model shipped, alongside Chinese chipmakers Cambricon and Hygon completing their own day-zero adaptations [7]. That groundwork is now scaling: DeepSeek is reportedly planning to deploy more than 160,000 Huawei Ascend accelerators at a data center in Inner Mongolia, a buildout Huawei itself appears to be preparing hardware for [8]. The company's trajectory marks a real shift - DeepSeek initially depended heavily on Nvidia processors, particularly the H800, before increasingly moving toward Chinese alternatives [9]. DeepSeek has publicly confirmed using open-source tools based on Huawei's Ascend chips, according to Reuters reporting [10].

The Skeptic's Case: Inference Milestone, Not a Proven Training Breakthrough

Not every reaction has been celebratory. Community discussion pushed back hard on the idea that DeepSeek has fully moved its training workloads onto Ascend hardware, pointing to earlier reporting that a prior DeepSeek training run on Huawei hardware reportedly failed and had to fall back to Nvidia for pre-training - meaning Ascend's role to date has been concentrated in inference serving, not the far more demanding training phase. That contested timeline is already playing out in public: one widely-shared X post claimed DeepSeek's next model, V5, would be the first trained fully on Ascend chips rather than Nvidia - a rumor not confirmed by any official statement, but one that shows how unsettled the training-versus-inference question still is. That distinction matters: CSIS's own framing is conditional - its warning about posing 'a much more significant threat to Nvidia' describes what would happen if DeepSeek fully committed its optimization techniques to Ascend and CANN, not a claim that it already has [5]. Technical reviewers raised a parallel caveat: a reported 99 percent 'hardware utilization' figure measures how efficiently the software uses Ascend's own performance ceiling, which is a separate question from how that ceiling compares to Nvidia's latest chips - and CUDA's real moat is less about peak throughput than decades of debuggers, profilers, and documentation that a new, open-source, API-compatible toolkit doesn't erase overnight just by existing. The honest read sits between the two extremes: real, substantive software and inference infrastructure progress, running well ahead of any confirmed training-scale replacement of Nvidia.

Historical Context

2025-03-07
CSIS analysis documented Huawei's Ascend 910C delivering roughly 60% of Nvidia H100 inference performance despite US export controls, framing the backdrop against which the later DeepSeek-Huawei software collaboration developed.
2025-09
TileLang-Ascend was first open-sourced, giving developers an initial domain-specific language for writing high-performance workloads on Huawei processors, ahead of the broader September 2026 six-component toolkit release.
2026-04-24
DeepSeek V4, a 1.6-trillion-parameter mixture-of-experts model, launched with 'day zero' adaptation support on Huawei Ascend 950PR/950DT chips, announced via livestream hours after the model's release; Cambricon and Hygon chips also completed day-zero adaptation.
2026-09-30
DeepSeek announced the open-sourcing of a six-component Ascend software stack (TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, DeepSelect) and disclosed the 128-chip Ascend 950 supernode collaboration and plans for a 160,000-plus accelerator deployment.

Power Map

Key Players
Subject

DeepSeek-Huawei Ascend Chip Software Partnership Challenges Nvidia CUDA

DE

DeepSeek

Chinese AI model developer that open-sourced the Ascend programming tools and is deploying Huawei chips at scale for inference workloads

HU

Huawei

Hardware provider of the Ascend 950-series NPUs; co-developed the compiler/kernel optimizations and the 128-chip supernode with DeepSeek

NV

Nvidia

Incumbent GPU/CUDA ecosystem provider positioned as the competitive target of the Ascend toolkit release

JE

Jensen Huang, CEO, Nvidia

Public commentator warning that DeepSeek standardizing on Huawei's tech stack poses a strategic risk to the US

HI

High-Flyer Capital Management

Quantitative hedge fund founded by DeepSeek founder Liang Wenfeng in 2015; the financial backer behind DeepSeek's origins

Fact Check

10 cited
  1. [1] DeepSeek Open-Sources Six-Part Ascend Software Stack
  2. [2] DeepSeek and Huawei Open-Source TileLang to Challenge Nvidia's CUDA
  3. [3] DeepSeek Open-Sources Infrastructure Components for Huawei Ascend Platform
  4. [4] DeepSeek's New Tools Pit Huawei Against Nvidia's CUDA
  5. [5] DeepSeek, Huawei, Export Controls, and the Future of the US-China AI Race
  6. [6] Nvidia's Jensen Huang Warns DeepSeek-on-Huawei Would Be a 'Horrible Outcome' for the US
  7. [7] Huawei Ascend, Cambricon, and Hygon Complete Day-0 Adaptation to DeepSeek V4
  8. [8] Huawei Prepares 160,000 Ascend 950DT Accelerators for DeepSeek Data Center
  9. [9] DeepSeek, Huawei Partner on Ascend AI Chip Programming Tools
  10. [10] China's DeepSeek Says It Used Open-Source Tools Based on Huawei Ascend Chips

Source Articles

Top 5

THE SIGNAL.

Analysts

“Says DeepSeek optimizing AI models for Huawei's Ascend chips instead of American hardware would be a serious strategic setback for the US, and that if future AI models are built on a different tech stack as AI diffuses globally, China could become superior to the US in AI.”

Jensen Huang, CEO, Nvidia
Warns against the shift to Huawei hardware

“Argues that if DeepSeek applied its CUDA-bypassing optimization techniques to strengthening Huawei's Ascend chips and CANN software ecosystem, it would pose a much greater threat to Nvidia than DeepSeek's model releases alone.”

CSIS (Gregory C. Allen)
Export controls are slowing but not stopping China's AI chip progress

“Says Huawei Ascend chips are China's best homegrown alternative to Nvidia, and that day-zero support for DeepSeek V4 shows top Chinese AI models can now run on domestic hardware.”

He Hui, Director of Semiconductor Research, Omdia
Independent industry analyst
The Crowd

“‼️ BREAKING: DeepSeek is building software for Huawei's AI chips to break its dependence on Nvidia, open-sourcing six core tools that take direct aim at CUDA. Huawei backed the project, which adapts TileLang, a coding language DeepSeek calls simpler than Nvidia's CUDA, for”

@@IntCyberDigest2604

“🚨 DeepSeek V5 Leak: Beats Astra >DeepSeek is reportedly preparing an imminent V5 launch >Founder Liang Wenfeng calls it the company's biggest bet yet >Rumored at 2 trillion parameters (not 3T) >Reportedly the first DeepSeek model to train fully on Huawei Ascend chips instead of”

@@Priyannkaaaa4572

“DeepSeek is preparing developers to move from NVIDIA to Huawei Ascend hardware. It released free, open-source software that lets developers keep familiar ways of working with NVIDIA but on Huawei hardware. Im thinking if this is their way to pave the way for a DGX Spark like”

@@plotarmordev216

“DeepSeek to order 160000 Huawei AI chips over Nvidia”

@u/CleaRSightZ594
Broadcast
DeepSeek Just Gave Huawei What Nvidia Fears Most.

DeepSeek Just Gave Huawei What Nvidia Fears Most.

DeepSeek + Huawei Just Went After Nvidia's CUDA Moat

DeepSeek + Huawei Just Went After Nvidia's CUDA Moat

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX)

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX)

DeepSeek-Huawei Ascend Chip Software Partnership Challenges Nvidia CUDA — AI News | Agentic Brew