CITIC Securities: Ascend 960 Released Ahead of Schedule, Entire Domestic Computing Power Supply Chain to Benefit
nashnova research
Huawei announced the Ascend 960 chip series will ship three quarters early, doubling the performance of its predecessor; CITIC Securities says domestic AI competition has shifted from single-chip specs to full-stack system-level rivalry, opening volume opportunities across the supply chain.
How far ahead of schedule is the Ascend 960, and how much faster is it?
The Ascend 960DT moves from late 2027 to Q1 2027 — three quarters early. The 960PR arrives in Q3 2027, one quarter ahead of plan.
FP8 compute reaches 2 PFLOPS, FP4 hits 4 PFLOPS — both double the prior generation. Max memory: 288 GB; memory bandwidth: 9.6 TB/s.
This means → Huawei's chip iteration pace is accelerating. The domestic supply chain can now deliver ahead of schedule, not just catch up.
What are "super nodes" and "optical interconnects," and why do they matter?
A super node — a massive compute unit linking thousands of AI chips with light signals instead of copper wires — can pack 4,096 cards in a single Ascend 960 super node, delivering up to 8 EFLOPS FP8 with 1 PB of memory.
Huawei also unveiled Hi-ONE, the industry's first mass-produced NPO optical engine — a near-package optical interconnect module that replaces electrical signals between chips with light, cutting latency and power.
In plain terms = a single chip can be as fast as it wants; if thousands of them can't talk to each other efficiently, the speed is wasted. Super nodes solve "how to make thousands of chips work as one machine," and optical interconnects are the high-speed link that makes it possible.
This reflects a pivotal shift: domestic AI competition is no longer about single-chip benchmarks — it is about system-level efficiency. Whoever connects more chips more tightly wins.
How much has been deployed, and where does commercialization stand?
The prior-generation Ascend 910C super node has over 1,000 units deployed. The Ascend 950 super node has entered volume commercial use.
Huawei's Lingqu UnifiedBus interconnect protocol — a communication standard for coordinating large numbers of chips — supports up to 1 million AI processors per cluster. Model floating-point utilization (MFU) on a 100,000-card cluster rose from 20% to 35%.
This means → going from 20% to 35% MFU delivers 75% more effective compute from the same hardware. The biggest cost in large-model training is not buying the chips — it is chips sitting idle. Higher utilization saves money directly.
What supporting hardware was launched alongside the super node?
Huawei released 11 key chips for the super-node cluster built on Lingqu, covering compute, interconnect, storage, and management functions.
It also launched OceanStor M900, an AI-oriented memory storage system. I/O latency and first-token latency drop more than 50% compared with traditional PCIe-plus-RoCE setups.
In plain terms = Huawei did not just build one main chip. It built every chip and storage module an entire "supercomputer" needs — that is what "full-stack in-house" means.
Which three investment themes does CITIC Securities highlight?
Theme 1: AI chip design companies — direct beneficiaries as the Ascend series scales.
Theme 2: Upstream supply chain — advanced process nodes, advanced packaging, advanced memory, and supporting components. Bigger super nodes mean more upstream demand.
Theme 3: Other related names — CITIC's report groups these as broad beneficiaries without naming specific stocks.
The report also flags risks: global macro weakness, trade-friction escalation, weaker-than-expected downstream demand, slower AI commercialization, slower domestic substitution, delayed wafer-fab expansion, and tighter U.S. sanctions on China.
What is the key test for whether this thesis plays out?
The core metric to watch: whether domestic AI compute can keep narrowing the system-level efficiency gap with leading overseas products.
This means → matching single-chip specs is only step one. Super-node MFU, optical-interconnect yield and cost, real-world large-model training throughput — these "system-level indicators" are the hard benchmarks for validating the supply-chain beneficiary thesis.
Put simply = Huawei's story is ambitious, but the market will ultimately judge one thing: when you wire thousands of your chips together, can your real training efficiency catch up with Nvidia's?
市场有风险,内容仅供研究参考,不构成投资建议。
