Morgan Stanley: China's AI Computing Power Shifting to System-Level Competition, Optical Interconnect Demand Accelerating

Alina Collins
Published todayAbout 16 min read

Morgan Stanley's latest report argues China's AI compute race is shifting from single-chip specs to system-level efficiency — whoever networks the most accelerators at low latency can offset the fab-process gap. This means → optical interconnects and supernode architectures become the next core value driver.

01

Why is the battleground moving from "single chip" to "system"?

China's wafer process remains constrained; single-chip specs cannot match Nvidia. This means → vendors are pivoting to assembling more accelerators into supernodes, using system architecture to close the single-chip gap.
WAIC 2026 sent a clear signal: fewer vendors launched standalone chips, while significantly more showcased rack-scale and cross-rack supernode designs.
Huawei, Moore Threads, Sugon, Metax, Alibaba, Enflame, and Biren are all expanding their Scale-up domains — the range of accelerators linked by high-speed direct connections — from 64–128 units rapidly toward 512, 640, 1,024, and beyond.
02

Copper inside the rack, optics across racks — how does the value of light scale?

Morgan Stanley's core call: short-distance intra-rack links stay copper — cheap and simple to integrate. But once the Scale-up domain spans multiple racks, copper's signal loss and power draw spike, and optical interconnects begin to penetrate.
In plain terms = short distances are fine on electrical cables; go longer and cables can't cope, so you switch to fiber — the bigger the supernode, the more optics you need.
Huawei's Atlas 950 SuperPoD already uses this hybrid approach: orthogonal cableless electrical interconnect for short range + optical links across racks, demonstrating 1,024 NPUs connected with a roadmap to 8,192, at roughly 2 TB/s per accelerator.
This reflects a structural shift: optical interconnects are no longer confined to the external data-center network — they are moving inside the supernode, and their value scales non-linearly with size.
03

What does the domestic Scale-up ecosystem look like?

The current landscape is dominated by proprietary vendor-specific interconnects, mirroring Nvidia's early NVLink path — vertical integration around each vendor's own chips first, open standards later.
Key configurations: Biren BLink 2.0 targeting 1,024 GPUs; Sugon scaleX640 supporting 640 accelerators per rack; Moore Threads MTLink 2.0 at 256 GPUs across two racks, ~800 GB/s per accelerator; Metax MetaXLink-E scaling to 128 units; Alibaba ALink + ICN at 128 per rack, ~800 GB/s.
Bandwidth comparison: AMD Helios ~3.6 TB/s, Nvidia GB300 NVL72 ~1.8 TB/s, Huawei ~2 TB/s, Moore Threads and Alibaba ~800 GB/s. Morgan Stanley cautions these figures cannot be ranked directly — accelerator definitions, link counts, protocol overhead, and topologies all differ.
04

NPO/CPO — from proof-of-concept to system deployment, how far along?

NPO (near-package optics) and CPO (co-packaged optics) deliver one core benefit: shortening the high-speed electrical path between the switch ASIC and the optical engine, boosting bandwidth density and cutting power. In plain terms = move the optical engine right next to the chip so signals travel less distance on copper — faster and more power-efficient.
Enflame and Lightelligence demonstrated an xPU-CPO prototype placing the optical engine beside the accelerator for direct optical output; at WAIC 2026, Enflame further showed an NPO-equipped GPU server targeting Scale-up domains of 512+ accelerators.
On the switch side, Lightelligence demonstrated 51.2T CPO and NPO switch modules. Morgan Stanley believes the design likely pairs an ASIC from Memory & Networking — which the report identifies as possibly Memory & Networking leader Memory & Networking firm Memory & Networking Asterfusion (盛科通信) — with Lightelligence's optical engine, incorporating silicon-photonic PICs and TIAs.
05

Can system-level TCO really undercut Nvidia?

Morgan Stanley modeled a 10 MW compute deployment: some Chinese AI chip solutions could deliver total cost of ownership (TCO) 30%–60% lower than Nvidia's China-facing processors; some leading domestic accelerators may match or beat Nvidia on per-token cost.
This means → even with weaker single-chip specs, the combined price, power, and inference-efficiency profile could let domestic solutions overtake Nvidia on TCO.
The report adds a caveat: these are model-based estimates assuming specific pricing, power, system configurations, and inference workloads — not all workloads will see the same result. Domestic chips still lag in software ecosystem maturity, model compatibility, and supply-chain scale.
06

What is the single most important variable ahead?

Morgan Stanley's structural call: Scale-up is the layer most likely to differentiate systems going forward; Scale-out remains the backbone for large-scale clusters.
How vendors organize hundreds of accelerators inside a supernode — proprietary vs. open, copper vs. optics, full-mesh vs. reconfigurable topology — directly determines model-parallel efficiency and system utilization. This is also the layer where Chinese vendors have the best shot at building differentiation and a self-sustaining ecosystem.
This reflects a deeper shift: the competition is no longer just "whose chip is faster" but "whose system is more efficient." The direction of Chinese cloud providers' capex will be the key proof point for whether this logic plays out.

Content is for reference only, not financial advice.

Morgan Stanley: China's AI Computing Power Shifting to System-Level Competition, Optical Interconnect Demand Accelerating · nashnova