A set of high performance kernels and communication libraries now allows developers to train models on Huawei Ascend hardware using a stack previously optimized for Nvidia. DeepSeek open-sourced this infrastructure—comprising TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect—to enabledomestic AI development in China. In official tests for Dense GEMM, the DeepGEMM Ascend library reached 99.8% of the theoretical hardware limit and 98% on MegaMoE.

The release was developed in collaboration with Huawei for the Ascend 950 128 card supernode, with deep optimizations for computation and communication. TileLang, a centerpiece of the suite, is a high-level language that permits writing operators in a Python-like format that compiles for Nvidia, AMD, and Ascend hardware. DeepSeek utilized TileLang to implement a large number of operators during the training of its V4 model series.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
  3. SOURCEthe-decoder.com
  4. SOURCE@zheanxu“achieved 99.8% of the hardware limit on GEMM and 98% on MegaMoE”x.com
  5. SOURCE@reuters“partners with Huawei to develop chip programming tools, reducing reliance on Nvidia”x.com
  6. SUPPORT@kyleichan“ordering 160,000 next-generation Ascend 950DT/960 chips, the company is building the infrastructure to move pre-training to Ascend by 2027”x.com
  7. SUPPORT@kyleichan“DeepSeek is gearing up to switch model training from Nvidia to Huawei chips”x.com
  8. SUPPORT@poezhao0605“shortages of components such as top-end memory will cap Huawei's output of its 950DT chip at the low hundreds of thousands this year”x.com
Markdown