Huawei, Sugon Launch 1024-Card Supernodes; Chinese Internet Giants to Scale Deployment by H2 2026
Chinese major internet firms plan to scale up deployment of supernode solutions starting in the second half of 2026. Huawei and Sugon have already unveiled 1024-card systems. Supernodes integrate multiple GPUs into a single logical supercomputer, boosting efficiency and reducing costs. Sugon's 100,000-card cluster at the National Supercomputing Internet Zhengzhou node achieves 90% utilization. The technology delivers 30% to 50% performance gains in PD-separated inference versus traditional setups.
A supernode is a system that deeply integrates multiple GPUs or accelerator cards from several servers into a logically single supercomputer using high-speed interconnect technology. It addresses bandwidth and latency bottlenecks in cross-machine communication for traditional clusters, improves GPU utilization, and reduces per-token costs. At the 2026 World Artificial Intelligence Conference, Huawei exhibited a live 1024-card Ascend 950 supernode comprising 20 cabinets, while Sugon showcased an immersion-liquid-cooled 640-card supernode. Baidu Tianchi, Alibaba Cloud Lingjun Zhenwu M890, Inspur Info Yuan Brain, and Muxi Xijing S600 were all placed in prominent positions on the exhibition floor. The entire supernode system costs more than an equivalent number of cards in a standard server, but offers higher compute density and a smaller footprint, enabling enterprises to accommodate more computing power in a limited space. DeepSeek, which had not previously used supernodes, has recently decided to adopt them and is currently negotiating pricing.
Domestically, three main approaches have emerged in the supernode sector. Huawei, Sugon, and Baidu are moving toward large-scale nodes. Huawei has already demonstrated a 1024-card supernode and will deliver 8192-card supernodes in volume in the fourth quarter of 2026. Its 384-card supernode, launched last year, has been commercially deployed in more than 750 units across industries including internet, finance, energy, education, and healthcare. Sugon uses immersion liquid cooling, with 640 accelerator cards per cabinet. A 100,000-card computing cluster built by Sugon has been deployed at the core node of the National Supercomputing Internet in Zhengzhou, achieving a utilization rate of 90%. Baidu Smart Cloud's Tianchi supernode is evolving from 256 cards to 512 cards, emphasizing a fully self-developed stack. Alibaba Cloud's Lingjun Zhenwu M890 uses a 64-card scale-up unit that can be expanded to larger clusters via scale-out and can be placed on the public cloud for on-demand use. Inspur Info launched a 64-card supernode last year and is introducing a 32-card product this year that can support trillion-parameter model inference. Muxi is based on a single-cabinet 64-card configuration, emphasizing CUDA ecosystem compatibility and horizontal scaling.
Tsingmicro uses a reconfigurable dataflow architecture that allows direct interconnection between chips. It has launched a 4096-card supernode with a peak computing power of 500 PFLOPS, and has deployed it in Beijing, Inner Mongolia, Xinjiang, Zhejiang, and other locations.
The ratio of online inference to training has approached 5:1 to 10:1, with most computing power used for inference. Although the entire supernode system costs more than a server with the same number of cards, it offers higher compute density and lower power consumption. Performance can be increased by 10 times while power consumption rises only 2 times. In the traditional model, the combined computing power of 10 machines is only equivalent to the performance of 6 machines. Supernodes minimize losses through high-speed interconnects, and in PD-separated inference scenarios, they deliver 30% to 50% better performance compared to machines with the same number of cards in a conventional setup. Supernodes involve multiple components including GPU, storage, communication, cooling, power supply, and software. They have now entered the stage of batch deployment, and various manufacturers have the capability to deliver complete systems.
In terms of ecosystem, Alibaba's Pingtouge and Muxi follow a CUDA-compatible path, while Huawei's Ascend and Kunlunxin pursue a self-developed ecosystem. Different choices exist in interconnect media and liquid cooling technologies. A consensus is that software and hardware architectures must be coordinated. The evolution of supernodes is closely tied to model technology roadmaps. Domestic model companies have split into two camps in the Attention approach: Linear Attention and Smart Attention. The evolution over the next two to three years will influence the hardware adaptation of supernodes.
Supernodes are becoming the foundation of AI infrastructure construction. The computing power of the United States is approximately nine times that of China. As more manufacturers gain the ability to deliver complete systems, the competitive focus is shifting to batch delivery, stable operation, and reducing per-token costs. Supply chain lock-in has become an important challenge.
Why this event matters
The event has a measured impact on 4 industrys. The strongest current signal is positive for Artificial Intelligence, with intensity 85/100 and 82% confidence over a medium term horizon.
Artificial Intelligence
- Direction
- positive
- Intensity
- 85
- Confidence
- 82%
- Horizon
- Medium term
Semiconductor Value Chain
- Direction
- positive
- Intensity
- 80
- Confidence
- 80%
- Horizon
- Medium term
Cloud Services & Data Centres
- Direction
- positive
- Intensity
- 78
- Confidence
- 75%
- Horizon
- Medium term
Network Equipment
- Direction
- positive
- Intensity
- 70
- Confidence
- 70%
- Horizon
- Medium term
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.