JPMorgan: AI Cooling Per-Rack Value Approaching $100,000

nashnova research
今天发布阅读约 14 分钟

J.P. Morgan estimates cooling-component value per rack rises from $66,300 for Vera Rubin to $98,500 for Vera Rubin Ultra — a 49% jump that signals thermal management is becoming one of the fastest-growing subsystems inside AI servers.

01

Why are cooling parts in one rack worth nearly $100,000?

J.P. Morgan breaks the $98,500 into three pieces: 18 compute trays ≈ $58,500, 9 switch trays ≈ $30,000, and rack manifolds ≈ $10,000.
Versus Vera Rubin, compute trays added roughly $18,000 and switch trays about $14,300; manifolds were largely flat.
This means → nearly all the incremental value sits *inside* the trays — the rack didn't get bigger; each tray just needs more cooling parts at higher specs.
For context: the GB300 rack was only about $48,300. In plain terms = across two product generations, cooling "content per rack" has roughly doubled — growing faster than the chip itself upgrades.
02

Why did cold plates jump from 2 to 6?

The driver is vertical power delivery (VPD) — a design that moves power-supply components to the back of the chip board. Once power devices, some memory, and capacitors sit on the backside, a single front-side cold plate can no longer handle the heat; the back needs its own plates.
In the VRU design, cold plates for two Bianca modules rise from 2 to 6, with 4 dedicated to the backside. Add the 3 plates already serving networking chips and the power-distribution board, and one compute tray involves 9 cold plates in total.
This means → dual-sided cold plates are not an optional upgrade but a hard requirement of the new chip architecture — skip them and the chip overheats.
Materials are also advancing: cold-plate zones contacting HBM — high-bandwidth memory, a type of stacked DRAM — may use diamond-copper composites, but only at specific contact areas, not as a full-plate replacement.
03

How much more do quick disconnects and lids cost now?

Quick disconnects (QDs) — detachable fittings in liquid-cooling lines: some VRU models upgrade to higher-spec ZQDs at roughly twice the average price, lifting per-tray QD value from $400–500 to about $750.
The bigger jump is in heat-spreader lids: VRU moves from a one-piece lid to a "removable cap + stiffener frame + screws" assembly, pushing per-GPU lid value to 4–5× the Rubin one-piece design.
This means → the lid is no longer just a metal cover; it now spans packaging protection, structural support, and the thermal path — three functions in one system component.
AMD and AWS show parallel trends: MI450 is expected to adopt a two-piece lid at roughly 10× the MI350 price; AWS stiffener frames run $10–20 each and will keep climbing. But base prices differ widely — "10×" cannot be read as a direct revenue comparison.
04

Lower-wattage ASICs — why is their cooling more expensive?

AWS Trainium 3 runs at 700 W TDP (thermal design power — the chip's maximum heat output), well below Rubin's 2,300 W. Yet its per-rack cooling value is about $86,000, higher than Rubin's $66,300.
The reason is structure: each compute tray packs 4 ASICs with 8 dual-sided cold plates plus a tray-sized base plate; the rack also holds 10 switch trays — one more than Rubin.
In plain terms = lower wattage ≠ cheaper cooling. More chips, more cold plates, and more switch trays can push total rack value above a higher-wattage alternative.
Google's TPU v8 takes yet another path: ~1,000 W chips, per-tray value close to Rubin, but the full rack uses only 32 QDs for a total QD value of about $1,800. This reflects a key point — more cold plates do not automatically mean more QD revenue; you have to break orders down to the component level.
05

When does mass production start, and what still needs proving?

J.P. Morgan expects some Rubin units to trial the new lid design in Q4 2026, with lid mass production pointing to Q1–Q2 2027; full VRU rack production is forecast for H2 2027.
This means → component trials, component mass production, and full-rack ramp are three distinct milestones — early lid shipments do not pull VRU rack volume forward.
The next one-to-two-year growth drivers point to switches using co-packaged optics (CPO) — integrating optical modules directly onto the switch chip — and near-package optics (NPO), plus 1.6T and 3.2T optical modules that could further raise liquid-cooling adoption.
But rising value ≠ realized profit. Dual-sided cold-plate qualification timelines, actual ZQD vs. MiniQD adoption, removable-lid ramp pace, and whether supplier gross margins can track the value uplift — these four checkpoints will determine whether the cooling sector's revenue potential converts into real earnings.

市场有风险,内容仅供研究参考,不构成投资建议。