Inside the Toolchain: Six Tools Built to Mirror CUDA
The September 30 release is not a single tool but a full stack: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, each ported to run on Huawei's Ascend chips [1]. Notably, all six correspond one-to-one with components DeepSeek had already open-sourced for Nvidia GPUs, meaning the company essentially rebuilt its entire Nvidia-facing training stack for a second hardware target rather than writing something new from scratch [3]. At the center is TileLang, pitched as a simpler, high-level programming language that lets developers write AI kernels without touching Ascend's lower-level instruction set directly [4]. DeepSeek and Huawei paired the software drop with a hardware milestone: a jointly built 128-chip supernode based on the Ascend 950, with every TileLang operator currently used in DeepSeek's own model training now carrying a high-performance Ascend implementation [5]. Early benchmarks suggest the abstraction layer isn't costing much performance - TileLang-Ascend GEMM workloads run at roughly 0.98x hand-written Ascend C code, vector operators around 0.96x, and cube-vector workloads around 0.95x [2]. DeepSeek framed the effort in explicitly ideological terms, stating that building an independent, self-controlled GPU software ecosystem starts with a general-purpose, easy-to-program high-level language capable of pushing hardware to its limits [5].
/dq/media/media_files/2026/09/30/deepseek-x-huawei-2026-09-30-17-31-38.png)


