Morgan Stanley: AI Memory Shortage to Persist for Years — Three Paths Forward Unlock Opportunities in Storage and Heterogeneous Computing
nashnova research
Morgan Stanley warns that DRAM shortages will persist through the current AI cycle — frontier models double in size every six months, yet fab construction takes years. The bank outlines three structural workarounds and reiterates overweight ratings on Micron, Astera Labs, and others.
Why will memory stay scarce for so long?
Frontier AI models double in size every six months; context windows grow 5–6× per year; inference concurrency keeps climbing — demand is expanding exponentially.
DRAM fabs take years from groundbreaking to production, leaving supply with almost no elasticity. This means → the shortage is not a blip but a structural mismatch between compute growth and memory supply.
Nvidia CEO Jensen Huang said publicly that the industry must address memory constraints through architectural innovation, not just wait for capacity. In plain terms = building more fabs alone can no longer keep up.
Path one: down-spec hardware to keep shipping?
Nvidia's Rubin architecture chose to cut memory specs: per-rack LPDDR5 dropped from a planned 54 TB to 28 TB; per-GPU HBM dropped from 288 GB to 192 GB, reducing stack height to protect shipment volumes.
This means → the bottleneck doesn't vanish — it migrates across the memory hierarchy. With less high-speed local memory, data such as KV caches (intermediate data stored temporarily during inference) gets pushed down to NAND flash.
The knock-on: cross-GPU data traffic rises, driving up network-interconnect bandwidth demand. In plain terms = faster networking and cheaper storage fill the gap left by scarce high-bandwidth memory.
Path two: split inference across purpose-built chips?
AI inference has two distinct phases: prefill (compute-heavy) and decode (memory-bandwidth-heavy). Running both on the same GPU wastes resources.
Heterogeneous inference — assigning each phase to different hardware — is gaining traction: general-purpose GPUs handle prefill; dedicated accelerators with large on-chip SRAM handle decode.
Key examples: Cerebras's wafer-scale engine and Groq's LPU architecture (acquired by Nvidia), both delivering far higher memory-bandwidth efficiency at the decode stage than conventional GPUs. This means → vendors specializing in decode have carved out a clear incremental market.
Path three: CXL unbundles memory from single machines?
CXL — Compute Express Link, a high-speed protocol that lets memory be expanded, shared, and pooled across processors instead of being locked to one chip.
Morgan Stanley estimates AI demand will push the CXL and related memory-attach chip market to roughly $6 billion by 2030, far exceeding the traditional CPU memory-expansion market.
This reflects a tiered memory architecture taking shape: hottest data stays in HBM, warm data moves to a CXL memory pool, cold data sinks to NAND. In plain terms = data is stored by "temperature," matching cost tiers to access frequency.
Where is Morgan Stanley placing its bets?
Storage: reiterates overweight on Micron (MU) and SanDisk (SNDK) — the down-spec is driven by supply scarcity, not weakening demand; the supply gap will keep absorbing capacity.
Interconnect: positions in CXL and scale-up interconnect leaders Astera Labs (ALAB) and Marvell (MRVL).
Heterogeneous inference: flags Cerebras (CBRS) and Nvidia (NVDA), which rounds out its heterogeneous stack via the Groq acquisition.
What could break this thesis?
AI compute demand growth falls short of expectations — model scaling slows, and the memory gap narrows on its own.
CXL adoption lags — the tiered architecture is delayed, pushing back the incremental case for interconnect names.
Memory capacity expands faster than assumed — if fabs ramp ahead of schedule, the supply gap closes early and storage valuations come under pressure.
市场有风险,内容仅供研究参考,不构成投资建议。
