行业资讯

The Computing Power Race Shifts to SuperPods: System-Level Integration Becomes Key to Cost Reduction and Efficiency Improvement in AI Infrastructure

2026-07-28 KLD Technology 0
image


In the second half of 2026, the focus of the computing power race is shifting from simply stacking more chips to system-level collaborative computing centered around **SuperPods**. This shift aims to solve the bandwidth and latency bottlenecks in cross-machine communication of traditional clusters by integrating a large number of GPUs/accelerator cards into a logically unified "supercomputer" through high-speed interconnect technology, thereby **significantly improving computing power utilization efficiency and significantly reducing the cost per token**.

Different technological approaches have emerged in the industry: the **"large node" approach** (such as Huawei's 1024 Ascend 950 cards and Sugon's 640 cards) aims to provide a massive memory pool for training trillion-parameter models; the **"small node" approach** (such as Alibaba Cloud and Inspur's 64/32-card solutions) emphasizes deployment economy and flexibility. Furthermore, vendors such as Tsingmicro are exploring a new architecture of **switchless, direct chip connection**.

The three main drivers of this surge are: a **surge in model parameters**, a **explosion in inference demand** (the online inference-to-training ratio has reached 5:1 to 10:1), and the **overall performance advantage brought by supernodes**—although the overall system cost is higher, by **reducing communication losses**, system performance can be improved by 30%-50% compared to ordinary clusters with the same number of GPUs.

**Interconnect Chips and Testing Challenges**

The core of achieving supernode integration lies in **high-speed interconnect chips and protocols**, such as Huawei's internal bus and the **ETH-X open protocol** promoted by ODCC. This places higher demands on **chip test sockets**: they must support the complete transmission of **ultra-high bandwidth signals** (such as 224G/448G SerDes), ensure **heat dissipation and signal integrity** under extreme power density, and adapt to new heat dissipation environments such as liquid cooling. As system complexity skyrockets, **high reliability and maintainability** also become critical, requiring test solutions to effectively ensure the stable operation of this expensive supernode system.

Supernodes mark a new stage in the AI ??computing power competition, entering a phase of **system-level engineering**. Through collaborative innovation of software and hardware, they are becoming a key foundation for reducing costs and increasing efficiency in AI infrastructure.