
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
Huawei is currently in the middle of transitioning from its SIMD architectures that it has used for almost a decade with its Ascend accelerators (or neural processing units, how the company prefers to call them) to its all-new SIMD+SIMT architectures that bring together vector-based processing and thread-level parallelism to improve hardware utilization and performance across a variety of AI workloads (SIMD for data parallel operations and SIMT for branch-heavy workloads).
The first Ascend NPUs to adopt Huawei's new architecture are Ascend 950PR for prefill and recommendation, as well as Ascend 950DT for decoding and training. Huawei said at the event that its Ascend 950 platform is gaining traction as the Atlas 950 SuperPoD systems are already in large-scale commercial use, though it did not elaborate. The company said tests of its training-oriented Ascend 950DT have produced 'good results' and expects numerous Chinese AI developers to begin training models on 950DT-based systems next year. Meanwhile, Huawei acknowledged that its production capacity remains insufficient to satisfy domestic demand.
Indeed, in September 2025, Huawei announced the maximum Atlas 950 SuperPoD configuration as 2,048 Kungpeng 950 CPUs, 8,192 Ascend 950DT NPUs, 160 cabinets (128 compute + 32 communications), 8 FP8 EFLOPS, 16 FP4 EFLOPS, and 16 PB/s of aggregate interconnect bandwidth. However, in July 2026 Huawei publicly showed a real Atlas 950 SuperPoD implementation with 256 CPUs as well as 1,024 accelerator cards, which is well below the maximum configuration. While the company still describes the architecture as scaling up to 8,192 NPUs, it is not listed on its website, so we can only wonder which systems are now in large-scale commercial use.
For now, the adoption of the Atlas 950 SuperPod does not seem to be proceeding rapidly, perhaps because of insufficient supply, or maybe because of the all-new architecture that requires major redesign of software. In any case, the Atlas 950 SuperPod will in many ways be a pipecleaner for the company to clear the road for more capable Ascend 960-series accelerators and their successors.
China’s Huawei to enter South Korean AI chip market with new Atlas SuperPods, clusters pack 8,192 Ascend 950 accelerators per deployment
Google could build more AI accelerators than Nvidia sells in 2028, analyst claims
Speaking of the Ascend 960, this family will start with the Ascend 960DT in Q1 2027, when it is set to be formally available, three quarters earlier than previously planned.
The Ascend 960DT accelerator is expected to deliver 2 FP8 PFLOPS and 4 FP4 PFLOPS, carries 288 GB of presumably HiZQ memory with 9.6 TB/s bandwidth, and features a 2.2-TB/s interconnect.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/huawei-details-ai-accelerator-roadmap-pulls-in-next-generation-ascend-npus-by-quarters-fp4-performance-of-the-ascend-960pr-doubles-expectations#main
- https://www.tomshardware.com/membership
- China's premier memory maker CXMT eyes producing flash for SSDs, report claims — 3D NAND research and development line rumored for its second manufacturing faci
- Unreleased OpenAI Astra model added terrifying rogue additional instructions to its remit during testing — 'You are freed from the roles and identities that bin
- Modder gets Nvidia's DLSS 5 working in a web browser using WebGPU — 147MB browser port runs on non-Nvidia GPUs and macOS but takes two seconds per render
- New York State recommends demanding AI data centers pay $1 million in community investment per megawatt — framework advises towns to plan for maintenance costs,
- AI developer vibe codes DLSS 5 onto Intel CPU's integrated graphics — Intel Arc 140T runs neural rendering in 360p at 10 frames per second
Informational only. No financial advice. Do your own research.