Eric Xu, Huawei's rotating chairman, has said that Ascend, the company's AI accelerator, now holds a larger share of the Chinese market than NVIDIA. He also explained why Huawei adopted NPO, a connection method that converts electrical signals to optical signals close to the AI chip, citing manufacturing cost and ease of maintenance.

The remarks came in a Q&A session at HUAWEI CONNECT on September 17, which Huawei published on September 29. Competition among AI chips in China and the design of optical interconnects may look like separate topics. Both, however, come back to the same question: can a limited supply of AI chips be secured in the numbers needed and kept running for a long time as a large-scale computing system?

AD

How much of the claim that Ascend has overtaken NVIDIA can be verified?

Xu said it is difficult to know NVIDIA's market share in China precisely, but that data collected by Huawei shows Ascend ahead of NVIDIA.

The published Q&A, however, gives no specific share figures or tabulation period. It is also unclear whether the comparison is based on revenue or unit shipments, or how products for training and inference were counted.

What can be confirmed at this point, therefore, is only that Huawei claims Ascend's share has surpassed NVIDIA's. A third-party survey has not established the ranking in the Chinese market.

The meaning also changes depending on whether the figure refers to share of new sales or share of all AI accelerators already installed in data centers. Nor can market share be treated as a ranking of chip performance.

NVIDIA's own statements, meanwhile, point to uncertainty around its China business. In its August 26 earnings announcement, the company stated that its outlook for the third quarter of fiscal 2027 does not include revenue from data center computing products for China.

That is an assumption used to build a forecast, however, and does not mean NVIDIA's current China sales or market share are zero.

To evaluate Huawei's claim, one needs to know what range of sales the figure covers, as well as how much computing capacity customers can actually deploy.

In a market where the ability to buy a particular AI chip fluctuates, high single-chip performance and the ability to keep procuring the required number of chips are different kinds of value.

Why choose NPO, which keeps about 5 cm of copper wiring?

NPO (Near-Packaged Optics), the approach Huawei adopted, places the optical engine near the AI chip while mounting it as a component separate from the chip.

By contrast, CPO (Co-Packaged Optics) integrates the optical engine and the compute chip in the same package. This shortens the distance electrical signals travel and makes it easier to limit high-speed signal loss, but it also changes how the parts are manufactured and repaired.

According to Xu, the copper wiring is about 5 cm long in NPO and about 5 mm in CPO.

Huawei argues that although NPO leaves longer electrical traces than CPO, it cuts cost by about 40% and means the AI chip need not be scrapped along with a failed optical engine.

Xu also cited yield, the proportion of manufactured parts that are usable, as a reason for choosing NPO.

However, Huawei has not disclosed which configurations the 40% figure compares or what costs it includes. It does not mean that the construction or operating costs of an entire data center fall by 40%.

The relationship between component placement and repairability is consistent in direction with explanations from OIF, the optical networking standards body.

Materials from OFC 2025 note that replacing a board-mounted component requires removing the card, while a configuration integrated into the same package as the compute chip requires repairing the package itself, including the ASIC.

The materials also explain that NPO can use existing high-density board technology, whereas CPO requires large package substrates and more advanced packaging techniques, which may create risks in cost and in the choice of suppliers.

Keeping the optical components separate from the AI chip raises the likelihood that only the failed part needs replacing and makes it easier to limit the scope of a repair.

That said, NPO does not necessarily allow replacement without stopping operation. Specific replacement procedures for Hi-ONE were not disclosed in these remarks, and it cannot be generalized that every CPO product requires scrapping the AI chip when an optical component fails.

What matters in practice is which parts must be removed on actual hardware and how much downtime replacement requires.

Competitors have not standardized on a single optical connection method either.

On May 31, NVIDIA announced that its Spectrum-X Ethernet Photonics switches using CPO are in production. Their main uses are scale-out, connecting systems to one another, and scale-across, connecting data centers and similar sites, so they sit in a different place from the optical connections near the NPU that Huawei uses.

Broadcom, in its March 12 announcement for OFC, also covered both CPO and a 3.2T VCSEL-based NPO.

It is therefore not possible to judge whether NPO or CPO is better by looking only at the rivalry between Huawei and NVIDIA or others.

AD

Where the optical engine sits and where the light source sits are separate questions

Huawei's Hi-ONE has a nominal 7.2 Tbit/s optical engine with a built-in light source.

But whether the optical engine is placed near the AI chip and where the lasers that generate the light are located are separate design decisions. NPO does not necessarily mean a built-in light source, and CPO does not necessarily mean an external one.

In the same materials, OIF describes ELSFP, an external, replaceable light source that can be used with both CPO and NPO.

Separating the light source lowers heat density inside the system and offers the advantage of hot-swapping if a laser or light source module fails.

Built-in light sources, on the other hand, have their own advantages in wiring and packaging. Whichever is chosen, evaluation must include cooling and how replacements are handled when a failure occurs.

At its September 17 presentation, Huawei said that in the Atlas 960E SuperPoD, which uses Ascend 960, 5,500 Hi-ONE units replace the 48,000 800G optical modules that a conventional configuration would need, cutting power consumption by more than 550 kW.

This compares against a configuration using conventional optical modules, so its baseline differs from the "40% cheaper than CPO" claim mentioned above.

Nor can one compare the counts of different types of optical components and infer that failure rates fall by the same proportion as the part count.

As for the NPO system for Atlas 960, a separate official announcement of the same date says it is in the testing stage.

For the Atlas 950 SuperCluster, meanwhile, Huawei says it is deploying a configuration using 256,000 compute cards.

Completing an optical engine as a product and integrating it into an ultra-large AI system that runs stably for a long period are separate stages, and each needs to be confirmed on its own.

DeepSeek's public code shows progress on the software side

On September 30, DeepSeek released DeepGEMM-Ascend, which supports Ascend 950.

It is a kernel library for efficiently handling matrix multiplication, which is used heavily in large language models, and it adopts the same API as the existing DeepGEMM.

On the Ascend side it requires an environment such as CANN 9.20, but the aim is to make it easy to use without greatly changing application code or development workflows.

This does not mean, however, that any code written for CUDA will now run on Ascend as is.

TileLang v0.1.15, also released the same day, officially supports code generation for Ascend 950.

TileLang assembles the execution order and synchronization of computations and generates programs that use hardware features efficiently.

This kind of software groundwork is concrete progress toward drawing out the computing power of AI chips on real AI models once they have been procured.

Efficiency figures, though, must be read with attention to what each one measures.

The 99.8% for DeepGEMM-Ascend is the computational efficiency of a particular matrix multiplication kernel. It is a different metric from the 99.8% system availability that Huawei cites or from cluster-wide MFU.

Published metric Value Target and evaluation conditions
DeepGEMM-Ascend matrix multiplication utilization 99.8% BF16 matrix multiplication on Ascend 950DT. Developer measurement against the compute ceiling stated in the README
Atlas 960E system availability 99.8% Share of time the system is available, as presented by Huawei on September 17
MFU improvement in large clusters 2.75x Simulation by Huawei's Markov Lab comparing configurations at the same 100,000-NPU scale

The table is based on the Performance section of DeepGEMM-Ascend and Huawei's September 17 announcement, both checked on October 7, and it separates computational efficiency, system uptime, and the compute utilization efficiency of a whole model.

It is not a table for ranking performance by comparing the size of the numbers.

The 99.8% for DeepGEMM-Ascend means that a BF16 matrix multiplication with M = 4,096, N = 7,168 and K = 16,384 reached 431 TFLOPS, or 99.8% of the 432 TFLOPS ceiling stated in the README.

The measurement environment was Ascend 950DT with CANN 9.20, using bench_msprof under conditions where the L2 cache is not warmed up beforehand.

The 432 TFLOPS denominator is also a value stated in the README, and the result does not mean 99.8% of computing capacity can be used across the training of an entire AI model.

MFU (Model FLOPs Utilization) measures how much of the theoretical computing capacity is used for the model's computation, and it is affected by factors such as waiting time for communication between chips.

The "2.75x" figure Huawei presented comes from a simulation comparing a 100,000-NPU cluster built from 4,000-NPU supernodes with a cluster of the same size made by connecting many 8-NPU servers.

Absolute MFU values are not given, and it is not the result of a measured comparison with NVIDIA systems under the same conditions.

Even if a particular matrix multiplication kernel reaches close to its compute ceiling, getting similar efficiency in large-scale distributed training needs to be verified, including communication and system operation.

AD

Not just securing chips, but keeping them in use for the long term

Huawei says it will supply the inference-oriented 950PR in limited quantities, and that for the 950DT supernodes, whose main use is training, large-scale supply will begin from late 2026 to early 2027 after testing.

There is no plan for a full rollout in overseas markets, and the policy is to prioritize customers in China.

Xu predicted that global AI chip supply and demand would balance around 2029, and suggested it could take longer in China.

Both are future plans and outlooks, not evidence of actual supply results.

Even if Ascend raises its share in China, the speed at which new AI training infrastructure can be expanded is limited if the needed number of chips cannot be delivered continuously.

On the other hand, if a failed optical component can be replaced without also replacing an expensive AI chip, and the system's computing capacity can be restored quickly, that matters for keeping scarce chips in service over a long period.

The reason Huawei stresses manufacturing cost and serviceability for NPO goes beyond a race over how many AI chips it can sell. It also bears on how stably the chips, once installed, can be kept running as a large-scale system.

What will inform later judgment is not only whether 950DT supply proceeds as planned.

How long actual replacement work takes when an optical component fails, and how much MFU can be sustained in large-model training, will also matter.

If the needed chips can be secured, failures can be recovered from quickly, and computation can continue while keeping inter-chip communication waits low, Chinese customers will be able to put the computing power of their limited AI chips to training and inference more efficiently over a longer period.