DeepSeek Completes the Ascend Software Stack

On September 30, DeepSeek open-sourced TileLang, a compute library, and a distributed communication library for Huawei Ascend, with the components corresponding to its open-source projects for NVIDIA platforms. This gives Ascend developers a more complete set of operator and cluster communication tools, but the ecosystem’s maturity will still depend on subsequent migration costs, performance data, and real-world deployments.
DeepSeek Brings a Suite of AI Infrastructure Components to Ascend
On September 30, DeepSeek announced the open sourcing of infrastructure components for Huawei Ascend's computing platform, covering the high-level language and compiler tool TileLang, computational libraries, and distributed communication libraries. According to the company, these components correspond one-to-one with the open-source components previously released for Nvidia platforms. In several key test cases, computational and communication performance has approached the hardware limits.
This is not the release of a new model, but the opening up of part of the "machine room" behind model training. Its practical significance does not lie in Ascend being able to replace Nvidia at no cost, but in whether developers can avoid writing another layer of hardware-adaptation code, with reusable implementations for foundational tasks such as operators, collective communication, and data processing.

Which Stages Are Covered, from Operators to Communication
This open-source release is not a single repository, but a group of components targeting different workloads. According to the official introduction, TileLang is responsible for describing computations through a higher-level programming approach, while the other components handle tasks including matrix operations, attention, data selection, vector computation, and cross-device communication.
| Component | Primary purpose | | --- | --- | | TileLang | High-level language and compiler tool for high-performance operators | | DeepGEMM Ascend | Accelerates matrix operations such as general matrix multiplication | | DeepEP Ascend | Provides large-scale cross-device communication capabilities | | TileKernels | Provides common vector-computation and data-access operators | | FlashMLA | Provides sparse attention operators for long-context processing | | DeepSelect | Implements efficient data selection |
These names do not refer to peripheral features. Matrix multiplication is the workhorse of Transformer computation, while attention implementations affect the efficiency of long-context inference and training. In scenarios such as mixture-of-experts models, data must be dynamically distributed across devices, and communication overhead can directly offset the gains from computational acceleration. Whether a model can run on a particular chip and whether it can be trained efficiently on that chip are not the same thing.
This release therefore appears to move the discussion from whether a chip is "usable" toward whether software can fully exploit it. Hardware specifications determine theoretical capability, while actual throughput depends on the compiler, operator libraries, communication stack, and coordination among them. Without these foundational components, model teams often have to complete the adaptation work themselves. The cost involves not only writing code, but also tuning, validation, and ongoing maintenance.
The Key to TileLang Is Not Just Writing Fewer Lines Than CUDA
DeepSeek describes TileLang as a simpler programming approach than CUDA, while also emphasizing that it can leverage chip-specific features to reach the hardware performance ceiling. These two statements point respectively to ease of use and performance. The real challenge is achieving both at once: if the language abstraction is too thin, developers still have to handle extensive low-level details; if it is too thick, it may be unable to precisely control data layout, parallel partitioning, and memory-access behavior.
It can be understood as an operator-oriented programming interface. Developers express computation and data flow in a more compact form, and the compiler then maps them to low-level instructions for the target hardware. For the Ascend version, DeepSeek says it encapsulates the underlying Ascend C instructions while retaining hardware performance, and that every TileLang operator used in training has a corresponding high-performance implementation on Ascend.
The word "corresponding" is important, but it also needs to be interpreted carefully. It means the project is attempting to provide reusable expressions and implementations for a class of operators across different hardware platforms. It does not mean that all code can run cross-platform without modification, still less that all models already deliver identical performance. Operator coverage, stability across changing tensor shapes, and the compilation and debugging experience must still be evaluated through repository code, test results, and real workloads.
DeepSeek also states that the TileLang approach was first validated on Nvidia platforms and has already supported the implementation of most operators used in training the DeepSeek V4 series. This provides important context: it is not a conceptual tool assembled temporarily for this release, but a programming approach with experience in real training scenarios. The next question is whether this experience can be fully reproduced on Ascend, and whether the community can add more operators and handle more edge cases beyond the official coverage.
Beyond Computation, Communication Is Another Constraint on Cluster Efficiency
For an operator running on a single card, optimization typically focuses on computational throughput, memory access, and parallelization strategies. In multi-card training, data exchange between devices also becomes a system bottleneck. DeepEP targets large-scale cross-device communication and directly addresses this problem. Even if every card in a cluster computes quickly, increasing the number of devices may not proportionally reduce training time if data distribution, synchronization, and communication scheduling cannot keep up.
The value of a communication library therefore cannot be summarized by a single bandwidth figure. It must work under specific parallelization strategies, device topologies, and workloads, while overlapping with computation as much as possible so that the chips still have useful work to perform while waiting for data. DeepSeek says that the relevant components approach the hardware limits in several key tests, but the currently available reference materials do not provide the test environment, metric definitions, comparison targets, or complete figures. For R&D teams, this is a positive official signal, but not a complete benchmark report that can be used directly for procurement or performance planning.
The remaining components fill different roles in the same chain. TileKernels covers routine vector computation and memory-access operations; FlashMLA targets sparse attention; and DeepSelect handles data selection. Such "supporting" operators often determine end-to-end performance: the fact that one stage is fast in a microbenchmark does not mean that the entire training or inference process will be fast. Data-format conversion, memory movement, and scheduling overhead can all consume the gains.
What This Means for the Ascend Ecosystem
For model teams using Ascend, the direct benefit is access to a set of foundational implementations drawn from large-scale model development, which can serve as a starting point for adaptation and optimization. In the past, teams had to manage interface and performance differences among hardware-vendor tools, general-purpose frameworks, and custom operators. Opening up the high-level programming tools and commonly used computational and communication components may reduce the pressure to repeatedly reinvent the same solutions.
For the chip ecosystem, competition in computing does not take place only at the level of peak chip specifications. Developers look at whether familiar frameworks are supported, whether core operators are complete, whether distributed training is stable, and whether problems can be diagnosed when they occur. When third-party teams invest in and publish reusable software, it can improve these conditions and give more developers an opportunity to evaluate the Ascend platform through real projects.
However, "one-to-one correspondence between components" should not be interpreted as meaning that the ecosystems are already aligned. The libraries, tools, tutorials, and developer experience accumulated over years on Nvidia platforms cannot be replicated through a single open-source release. Operator coverage, performance stability, debugging tools, and framework integration across platforms all require continued development. Whether open-source repositories have active maintenance, responsive issue handling, and version-compatibility strategies will likewise determine whether they can evolve from a one-time release into production dependencies.
A more realistic assessment is that this open-source release lowers some of the software barriers to using the Ascend platform and increases the possibility of cross-hardware adaptation for model teams. It has not demonstrated that all workloads can be migrated smoothly, nor has it provided enough public data to compare the cost of full-system training. Observers should distinguish performance claims from reproducible benchmarks.
Collaboration Extends from Code to Supernode Development
DeepSeek says that Huawei's team provided support during development for the Ascend platform, and that the two sides jointly advanced a 128-card supernode solution based on Ascend 950, with deep optimization of computation and communication. This indicates that the effort behind the open-source release involved more than retargeting existing code to another architecture; it also included hardware-software co-adaptation.
The focus of a supernode solution is to enable multiple accelerators to work together in a more tightly integrated system. For training workloads, the arrangement of computation and communication needs to be designed jointly. If inter-device communication paths, memory hierarchies, and operator implementations are disconnected from one another, optimizing a single library cannot guarantee end-to-end cluster efficiency. Whether the collaboration between DeepSeek and Huawei can turn these optimizations into open, maintainable components will be more important to observe over the long term than the performance statements made on the day of release.
This also provides a more concrete perspective on "domestic computing-power substitution." Substitution is not simply a matter of whether a chip can start a model. The question is whether, after migration, it can maintain acceptable development efficiency, training stability, and total cost. Open software components at least allow more teams to participate in validation instead of relying solely on adaptation results produced internally by vendors and a small number of large-model teams.
What Developers Should Watch Next
For infrastructure teams, the first steps should be to examine the hardware and software version requirements, build procedures, operator coverage, and test scripts in the repositories, then rerun the tests with their own model shapes and parallel configurations. It is particularly important to distinguish microbenchmarks from end-to-end results: the former help identify the ceiling of an individual operator, while the latter reflect training throughput, scaling efficiency, and actual engineering costs.
For framework and compiler developers, TileLang's abstraction boundaries and its mapping to Ascend C are areas worth studying. Whether it can make common operator implementations easier to read and maintain while still allowing fine-grained tuning on critical paths will affect whether more teams adopt the tool. The community will also need to pay attention to error diagnostics, performance analysis, documentation completeness, and version compatibility, rather than merely whether examples compile successfully.
For project decision-makers, these components can serve in the short term as a foundation for evaluating the Ascend platform and conducting prototype validation. It would be unwise to estimate the cost of an entire training system solely from the claim that performance is "close to the hardware limit." Testing should take into account the target model, cluster size, network topology, software-stack versions, and operational capabilities. Bottlenecks can differ completely across workloads: strong matrix-operator performance does not mean that communication-intensive tasks will benefit to the same extent.
The real significance of this open-source release is that it moves the discussion from chip specifications toward software implementations that developers can inspect. Whether DeepSeek can extend to Ascend the operator-development approach validated on Nvidia platforms, and whether these components can attract external contributions, receive sustained maintenance, and reproduce their performance across more real-world tasks, will determine the release's long-term significance. For the AI infrastructure industry, that matters far more than simply adding a few more repositories.
Open-Source Projects
- TileLang: High-level language and operator compiler tool.
- DeepGEMM Ascend: Matrix computation component for the Ascend platform.
- DeepEP Ascend: Distributed communication component for the Ascend platform.
- TileKernels: Vector-computation and memory-access operators.
- FlashMLA: Sparse attention operators.
- DeepSelect: Data selection component.
References
- ITHome: DeepSeek Open-Sources Infrastructure Components for Huawei's Ascend Computing Platform: Summarizes the scope of the open-source release, component purposes, and information about the Ascend version of TileLang.
- DeepSeek's Official Related Open-Source Projects: View DeepSeek's publicly available code repositories and project updates.



