Six Tools, One Goal: Replacing CUDA's Developer Experience on Ascend
The release is not one tool but six: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, covering matrix math, distributed communication, attention processing, and top-K selection for Huawei's Ascend NPUs [1]. The centerpiece is TileLang, pitched explicitly as a simpler way to write 'kernels' - the small, performance-critical programs that run directly on chip hardware - without requiring the deep hardware expertise CUDA historically demanded [2]. DeepSeek's own framing is pointed: 'To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential.' On the performance side, industry coverage citing GitHub reporting claims FlashMLA reached roughly 95 percent of theoretical performance on Ascend 950 for some processes [3]- a notable efficiency claim, though one that measures software optimization rather than a head-to-head comparison against Nvidia's newest silicon. The toolkit also underpins a jointly engineered 128-chip Ascend 950 'supernode,' tuned by DeepSeek and Huawei together to balance computation and data movement across the cluster [4].



