China's AI Supernode Race Shifts Toward System Integration and Optical Interconnects

Taylor Wilson
Published todayAbout 14 min read

At WAIC 2026, Huawei, Sugon, ZTE and others unveiled AI supernode products, signaling the competition has moved from single-chip performance to system-level orchestration — whoever can keep thousands of accelerators running together reliably now holds the edge.

01

What are supernodes actually competing on?

A supernode — a cluster linking hundreds or thousands of AI accelerators into one system — is no longer judged on how fast a single chip runs. The real test is whether processors, networking, memory, cooling, and software can work in concert.
Huawei publicly showed its Ascend 950 supernode for the first time: 1,024 NPUs, 1 EFLOPS at FP8, 256 TB unified memory, interconnect latency of just 3 microseconds, targeting trillion-parameter model training.
This means → the dividing line in Chinese AI hardware has shifted to system orchestration. A chip with high benchmarks that cannot scale across a cluster is irrelevant in this round.
02

Why has the bottleneck moved from chips to communication?

Model parameters are approaching the trillions and context windows have stretched from 256K to over one million tokens. Computation must span thousands of accelerators — the speed at which chips talk to each other is now the binding constraint.
Biren's BLink 2.0 targets congestion control and link recovery. Baidu's XPU-Link pairs custom interconnects with programmable switches, supporting 32 to 1,024 cards in a single high-bandwidth, low-latency domain.
In plain terms = the old bottleneck was "not computing fast enough." The new bottleneck is "not moving data between chips fast enough" — communication drags down the entire pipeline.
03

How can one dead chip paralyze a system worth over ¥100 million?

A single supernode can draw 400 kW; one high-density rack hits 48 kW — far above the roughly 6 kW typical of Chinese data-center racks a few years ago. Power delivery, cooling, cabling, and fault recovery are now key differentiators.
One accelerator failure can halt a system valued at over ¥100 million. Vendors must handle fault detection, state checkpointing, and workload redistribution.
This reflects a deeper truth: the real barrier in the supernode race is not peak compute — it is reliability. Running fast and running stable are two different things.
04

Why is Alibaba Cloud bringing supernodes to the public cloud?

Alibaba Cloud launched Lingji Zhenwu M890 — a 64-card instance with 800 GB/s inter-card bandwidth, delivering three times the training throughput of the prior Zhenwu 810E. It supports inference for mixture-of-experts models with tens of trillions of parameters and is now in invite-only beta in Ulanqab.
The Lingji platform supports 130,000 heterogeneous cards per cluster, scaling toward one million cards, with 99.7% average availability and minute-level fault recovery.
This means → supernodes used to require self-built clusters that only the largest companies could afford. Alibaba Cloud is turning them into an on-demand public-cloud service, putting trillion-parameter training resources within reach of smaller teams.
05

Fully integrated vs. open architecture — which path works?

China's supernode market is splitting into three lanes: Huawei's full-stack proprietary approach; ZTE and H3C leaning toward open compatibility; and Enflame, Biren, and MetaX pairing domestic accelerators with open joint architectures.
ZTE and Enflame jointly launched the Yunsui ESL64-O with a backplane-free OEX design, zero internal cabling, lower interconnect costs, and support for multiple domestic AI chips.
Huatai Securities projects China's supernode market could reach ¥341.4 billion (roughly $50.4 billion) by 2028, a compound annual growth rate of about 194% from 2026. This reflects that capital markets already view supernodes as the next major lane in AI infrastructure.
06

Why are optical interconnects and advanced packaging moving center stage now?

Copper signals degrade noticeably beyond about 3 meters; optical links can stretch hundreds of meters. Electrical interconnects are hitting physical limits, and the shift to optics is accelerating. Biren unveiled a 1,024-card near-package optics (NPO) supernode that places optical engines next to GPU modules, eliminating power-hungry digital signal processors.
Enflame and Xianfeng jointly launched a CoPoS — chip-on-panel-on-substrate — packaging solution using panel-level and glass-related technologies to improve planarity and high-frequency signal performance, offering another route around process-node constraints.
In plain terms = chips "talking over copper wires" is reaching its ceiling. Switching to light for data transmission is the most practical breakthrough path today; advanced packaging lets chips boost performance without shrinking to a smaller process node.

Content is for reference only, not financial advice.

China's AI Supernode Race Shifts Toward System Integration and Optical Interconnects · nashnova