On September 30, 2026, DeepSeek released a language foundation for developing compute programs, along with compute and communication libraries, for Huawei's AI chip, Ascend.
The effort adds Ascend support to a software suite DeepSeek had built up for NVIDIA GPUs, keeping the APIs and development workflow that developers use as consistent as possible. Huawei has also supported the development, and benchmarks for standalone matrix multiplication came close to Ascend's theoretical peak performance.
However, a shared API does not mean that supported features or data layouts are identical. Some of the multi-chip communication results were also measured in a development environment that predates general availability.
Below, we look at what becomes common when moving from an NVIDIA-centric development environment to Ascend, and which parts still require hardware-specific optimization.
Six software components cover everything from kernel development to inter-chip communication
TileLang itself has gained an official backend for Ascend 950. A backend is the mechanism that converts code written in TileLang into a form that can run on the target chip.
The official changelog distinguishes between the external Ascend adapter released on September 29, 2025, and the support now integrated into TileLang itself.
In other words, this is not the first time Ascend support has appeared. What has changed is the chip generation supported and the position of that support within TileLang.
TileLang is a dedicated language for developing "kernels" using Python-like syntax. Here, a kernel means an individual compute program that runs on an AI chip, such as matrix multiplication or an attention mechanism.
The main software released or updated this time breaks down by role as follows.
| Component | Processing it handles | What was released or updated for Ascend |
|---|---|---|
| TileLang | Writing and compiling kernels | Code generation for Ascend 950, coordination of execution order and synchronization |
| DeepGEMM-Ascend | Matrix multiplication at the core of AI computation | BF16, FP8 and FP4 operations, processing for MoE |
| DeepEP-Ascend | Communication between multiple chips | Dispatching data to experts and aggregating results |
| TileKernels | Surrounding operations and data processing | Dozens of kernels, including quantization and MoE routing |
| FlashMLA | Attention mechanism | Sparse attention that references only the necessary tokens |
| DeepSelect | Candidate selection | TopK processing that picks the highest scores |
Together, these correspond to the series of operations needed to actually run an AI model.
In MoE (Mixture of Experts), a subset of experts is chosen to handle each token, and the data is sent to the other chips where those experts are located.
Matrix operations are executed there, and the results are sent back to the original location and aggregated.
Speeding up matrix operations alone is therefore not enough: if waiting time arises in inter-chip data transfer or aggregation, the chip's performance cannot be fully used. Both faster computation through DeepGEMM and faster communication through DeepEP are needed.
Migrating from CUDA means less manual work, but per-chip optimization remains
DeepGEMM-Ascend is described as fully API-compatible with the existing DeepGEMM, allowing the same package name and development workflow to be used.
TileKernels also uses the same Python API and selects the appropriate backend for the runtime environment.
In DeepEP-Ascend, the public buffer API has been aligned with the NVIDIA version, and the design lets training and inference share a common interface.
The main aim of this standardization, then, is to let model code avoid major rewrites at the points where it calls the libraries.
Inside, however, processing is needed to absorb the differences between hardware.
According to the TileLang guide for Ascend 950, matrix operations are handled by the Cube core and surrounding vector operations by the Vector core.
When a developer describes the computation and data transfer, the compiler analyzes how to assign work to each core and the dependencies between them, then adjusts execution order and synchronization.
TileLang takes over some of the buffer management and synchronization that developers would need to specify in detail when using Ascend C directly.
Still, developers must specify things such as the data types to process and the units into which computations are split.
The correction factors used in DeepGEMM's low-precision operations are also laid out in memory in a different format on Ascend than on NVIDIA GPUs.
Even when the API can be shared, the internal data representations of the hardware are not unified.
DeepJIT standardizes, across both platforms, the procedure of compiling needed kernels at runtime and saving and reusing the generated binaries.
However, differences between hardware still remain in kernel source code and the settings passed to the compiler.
These are therefore not mechanisms that automatically convert CUDA code to Ascend as is.
The approach is to share the upper-level APIs and development procedures as much as possible, while preparing separate implementations for NVIDIA and Ascend underneath.
A concrete example can be seen in the FlashMLA technical report published by DeepSeek.
In the kernel that handles sparse attention, one Cube core and two Vector cores are combined, dividing up matrix operations, data transfer, and the conversion of low-precision data into a computable format.
The KV cache held during inference and the correction factors are also placed adjacent to each other in memory, reducing the work of copying them separately.
Porting to a different AI chip requires more than rewriting the same computation with different instructions. It means redesigning how data is laid out and moved internally to suit the chip's structure.
The 99.8% figure is for standalone matrix operations approaching theoretical performance
In the benchmark published for DeepGEMM-Ascend, dense matrix multiplication in BF16 reached 431 TFLOPS.
This corresponds to 99.8% of the published hardware ceiling of 432 TFLOPS.
The measurement used Ascend 950DT and CANN 9.20, and was run with the target data not resident in the L2 cache.
The matrix size was M=4096, N=7168, K=16384. For FP8-by-FP8 operations at the same size, the result was 861 TFLOPS, or 99.5% of the 865 TFLOPS ceiling.
Even if an AI chip has high theoretical compute performance, waiting time arises if data cannot be supplied to the compute units fast enough.
The result indicates that, at least for the measured matrix sizes and data precisions, the Ascend-optimized implementation kept the compute units running close to their upper limit.
This suggests that the design of maintaining a common API while optimizing internally for the hardware led to high performance in specific matrix operations.
However, the 99.8% figure does not mean that the full chip performance of 99.8% can be extracted across entire model training or inference services.
Real models combine many kinds of processing besides matrix multiplication, including attention, inter-chip communication, and operations on small amounts of data.
This figure alone cannot be used to judge overall performance differences from NVIDIA GPUs or to compare the cost of a single inference.
How close a standalone compute kernel gets to theoretical performance and how much speed is obtained when running an entire model need to be evaluated separately.
Scaling to 128-way parallelism reveals communication performance challenges
DeepEP-Ascend also publishes communication performance as the scale of expert parallelism is varied from 8 to 128.
Each process participating in distributed processing is called a "rank," and the published table shows the range of bandwidth measured at each rank.
The measurements used Ascend 950DT and CANN 9.2.0, in a Clos network environment connecting multiple compute nodes.
| Expert-parallel participants | Dispatch bandwidth (GB/s) | Combine bandwidth (GB/s) |
|---|---|---|
| EP8 | 373–375 | 345–347 |
| EP128 | 313–320 | 272–278 |
The source is the Performance table in DeepEP-Ascend.
The settings were a token capacity of 16,384 per rank, a hidden dimension of 7168, and 6 experts selected out of 256 for processing.
FP8 is used for data dispatch and BF16 for aggregating results.
The measured time includes everything from the start of communication to waiting for completion, but excludes the final post-processing. In addition, the measurement environment used a test HDK that was manually configured with extra settings.
A simple calculation using the midpoints of the ranges DeepSeek published shows that scaling from EP8 to EP128 lowers per-rank bandwidth by about 15.4% for dispatch and about 20.5% for combine.
This was calculated as (EP8 − EP128) ÷ EP8 × 100, taking dispatch as falling from 374 GB/s to 316.5 GB/s and combine from 346 GB/s to 275 GB/s.
However, this figure was computed by our editorial staff from the midpoints of the published bandwidth ranges and is not an average measured value.
Nor does it mean that total system-wide communication volume, or whole-model processing speed, falls by the same proportion as the number of participating ranks increases.
According to DeepSeek, for dispatch at EP32 and below, performance reaches about 90–95% of the bandwidth the physical network can actually deliver.
Optimization continues for larger parallel configurations and for the combine step.
Combine requires adding up the results arriving from each chip, and conflicts also arise because both AI computation and communication use high-bandwidth memory bandwidth.
Once matrix operations are accelerated close to the hardware's limit, performance depends not only on how fast data is sent between chips but also on how efficiently received data can be aggregated.
Conditions remain before the released code can be used in production environments
The communication performance published for DeepEP-Ascend was measured in an environment where a PoC HDK provided for DeepSeek's testing was manually configured with additional settings.
HDK refers to the software environment needed for hardware development, including drivers and firmware.
This configuration is not generally distributed at present.
According to the DeepEP-Ascend documentation, a commercial HDK for Atlas 850E containing the required settings is scheduled for release in mid-October 2026, around the 15th.
This is a release schedule stated as Huawei's, and the benchmark is not a result reproduced in a generally available commercial environment.
Feature differences also remain between the CUDA and Ascend versions.
The Ascend version of FlashMLA supports prefill and incremental decoding for sparse attention, but some fused kernels that run multiple operations together, and some kernels for dense attention, are supported only in the CUDA version.
DeepSelect for Ascend is likewise limited to BF16 input, so the FP32 sampling processing supported by the CUDA version is not equally available.
In addition, this version of FlashMLA drops support for Hopper and some older models, and the KV cache format has been changed. When deploying into an existing environment, it is therefore necessary to check the version used and its compatibility.
The communication features needed for distributed training are also not all complete yet.
DeepEP-Ascend says kernels for "all-reduce," which sums values across all ranks, and "reduce-scatter," which distributes the aggregated result across multiple ranks, are currently under implementation.
Some communication for balancing load among experts also remains unimplemented.
The release of training-oriented APIs should therefore be considered separately from whether large-scale AI model training can be completed from start to finish on Ascend alone.
Even so, when the model developers themselves release the compute and communication processing they need in open form, other developers can study the hardware-specific implementations and reuse them for other models and services.
The collaboration between DeepSeek and Huawei could evolve the model-specific optimizations for running DeepSeek models fast on Ascend into a broader AI development foundation.
Going forward, what matters is whether similar performance can be reproduced on the commercial HDK once it is generally available, and how far the unsupported communication and compute features get implemented.
If measurements with real models and services can further confirm inference latency, throughput, and training performance, developers comparing NVIDIA GPUs with Ascend will find it easier to judge based not only on raw compute performance but also on the effort required for porting and on performance in actual operation.
