On September 17, 2026, Huawei significantly overhauled its rollout plan for the "Ascend 960" AI accelerator. The training-focused Ascend 960DT will ship in Q1 2027, while the inference-focused 960PR will arrive in Q3 of the same year. The previous year's plan had called for a single, undifferentiated Ascend 960 in Q4 2027.
What stands out is the acceleration of up to nine months—but the change goes beyond dates. Huawei has made clear its strategy of splitting compute chips by use case while co-designing optical connectivity, memory, and storage as an integrated whole. Measuring the competition with NVIDIA purely by single-chip peak compute performance misses this intent.
The "Two-Track" Split Matters More Than the Nine-Month Acceleration
Comparing the official roadmaps from 2025 and 2026 makes the shift in product planning clear.
| Announcement Date | Product | Primary Use | Planned Timing | Evidence Status |
|---|---|---|---|---|
| September 2025 | Ascend 960 | Announced without training/inference split | Q4 2027 | Company plan |
| September 2026 | Ascend 960DT | Training | Q1 2027 | Company plan, moved up 3 quarters |
| September 2026 | Ascend 960PR | Inference | Q3 2027 | Company plan, moved up 1 quarter |
| September 2026 | Ascend 970 | Not disclosed | 2028 | Only year disclosed |
| September 2026 | Ascend 980 | Not disclosed | 2029 | Only year disclosed |
The Ascend 960, slated for Q4 2027 in the 2025 announcement, split in 2026 into the training-oriented 960DT (Q1 2027) and inference-oriented 960PR (Q3 2027). The acceleration and the use-case split happened simultaneously.
While Huawei has divided the 960 generation by use case, it has not explained the detailed design differences between the two products. Moreover, the 2026 official text does not disclose detailed compute performance, memory capacity, or manufacturing process specifications for either the 960DT or 960PR. Some media reports have cited detailed figures based on event projection materials, but it would be premature to treat numbers that cannot be reconfirmed in official published documents as finalized specifications.
Huawei states it will launch the Ascend 970 in 2028 and the 980 in 2029, updating one generation per year. It says compute specifications will double under what it calls the "Tau Scaling Law," with memory bandwidth, capacity, and interconnect bandwidth also increasing. However, this is a technical target set by Huawei, not an independently verified result.
The Unit of Competition Shifts From Chip to SuperPoD
In large-scale AI, communication latency becomes increasingly impactful as more accelerators are added. According to Huawei, in a 100,000-NPU cluster built from conventional servers, communication accounts for a significant share of training time. The company's Markov Labs simulated a configuration using 8-NPU servers versus one using a 4,000-NPU SuperPoD, finding the latter achieved 2.75 times higher Model FLOPs Utilization (MFU). While this isn't a real-hardware benchmark, it illustrates the problem Huawei is trying to solve.
The foundation for this is the Peerium Computing Architecture and UnifiedBus. Huawei combines nested parallel processing, unified memory addressing across physical nodes, and peer-to-peer connections. UnifiedBus is designed not just to link NPUs together, but to connect CPUs, memory, SSDs, and networking equipment under a single protocol.
The goal isn't simply to compensate for weaker chips through sheer numbers. By operating many NPUs as a single logical computer, Huawei aims to reduce idle compute time. Even with identical peak performance, less time lost to communication delays means higher effective performance for training and inference.
Proximity-Mounted Optics Target Power and Failure Points, Not Just Bandwidth
The new Atlas 960E SuperPoD accommodates up to 4,096 NPUs, with Huawei-stated figures of 8 EFLOPS FP8 compute performance and up to 1PB of HBM capacity. But the key technical shift this time lies not in the number of compute units but in the connection method. Huawei adopted proximity-mounted optics, placing optical engines close to the compute units.
The company's Hi-ONE transmits 7.2Tbit/s per unit. The Atlas 960E uses 5,500 of these units, eliminating the need for what would otherwise have been 48,000 800G optical modules. Huawei claims this substitution cuts power consumption by more than 550kW, doubles mean time between failures, and achieves 99.8% system availability.
The denominators behind these figures require caution. The 550kW-plus reduction is an absolute figure tied to changing the optical connection method—it is neither the total power consumption of the entire Atlas 960E nor a reduction rate. As for the 99.8% figure, the test period, the availability baseline used for comparison, and what was counted as a failure have not been disclosed. Verifying the true value of proximity-mounted optics requires measuring power consumption and repair times under identical conditions in actual operation.
Has the 15,488-NPU Version Disappeared?
The Atlas 960 SuperPoD that Huawei unveiled in 2025 was a massive configuration housing up to 15,488 Ascend 960 units. Stated figures included 30 EFLOPS FP8, 60 EFLOPS FP4, 4,460TB of memory, and 34PB/s interconnect bandwidth. It was slated for a Q4 2027 launch.
By contrast, the Atlas 960E announced in 2026 houses 4,096 NPUs. Comparing counts alone, it's smaller—but Huawei's public materials don't clarify whether the 4,096-NPU version replaces the 15,488-NPU configuration or exists alongside it as a separate option. The memory figures also differ in framing: the earlier figure represents total capacity, while the latter is expressed as "up to 1PB of HBM"—making it unclear whether the two use the same denominator.
The products are also at different stages of maturity. As of September 2026, Huawei describes the 256,000-card Atlas 950 SuperCluster as being in deployment, while the proximity-mounted optics version of the Atlas 960 is described as still in testing. The figure of up to 1 million NPUs likewise refers to an architectural capability for linking multiple SuperPoDs together, not a track record from an operational system.
Software and Supply Remain the Barriers
Even if hardware can be connected at massive scale, adoption won't spread unless existing models can run on it without friction. In September 2026, the PyTorch Foundation reported that the Accelerator Integration Working Group—co-chaired by Huawei and Intel—is advancing hardware onboarding guidelines, continuous integration across multiple repositories, and device-agnostic testing. It's confirmed that Huawei is participating in efforts to expand PyTorch's multi-backend support. But this doesn't mean Ascend itself is fully integrated, nor that it has accumulated the same software assets and operational experience as CUDA.
Another barrier is more physical than the roadmap suggests. According to Reuters, Huawei's rotating chairman Eric Xu said at a press conference that the company lacks sufficient supply capacity even to meet domestic demand within China, and has no plans for a full-scale overseas rollout. Moving up the launch timeline by nine months won't increase the total compute available to customers if HBM supply and manufacturing/packaging capacity can't keep pace.
The competitiveness of the Ascend 960 won't be determined simply by adding product names to a schedule. Only once mass-production volumes for the 960DT and 960PR, independently verified training and inference performance, real-world measurements of proximity-mounted optics' power consumption and availability, and the engineering effort required to port major models are all in place can it be judged that Huawei's SuperPoD strategy has moved from stated figures to practical reality.
