Nvidia Testing Plans to Cut HBM Capacity on Rubin Ultra Chips

0xBroomberg
Published todayAbout 9 min read

Nvidia is testing multiple Rubin Ultra GPU prototypes, some with memory as low as 192GB — down from the announced 1TB — partly because it may not secure enough advanced HBM chips. This means → the flagship next-gen AI chip's final spec remains unfixed, with memory supply reshaping the design itself.

01

1TB down to 192GB — how deep is the cut?

CEO Jensen Huang unveiled Rubin Ultra at this year's developer conference: 1TB of HBM4E memory across 16 memory stacks per GPU.
But three people familiar with the matter say Nvidia is now testing at least three prototype versions with fewer stack layers, lower capacity, and some using the older HBM4 spec instead of HBM4E.
The lowest-memory prototype holds just 192GB; another sits at 256GB. In plain terms = the most aggressive cut strips out more than 80% of the originally announced memory.
For context, Nvidia's current-generation Vera Rubin chip tops out at 288GB of HBM4 — meaning some Rubin Ultra test builds carry less memory than the chip already shipping.
02

Why cut the memory?

The core reason: Nvidia may not be able to source enough advanced HBM chips to support the original design. HBM — high-bandwidth memory, a technology that stacks multiple memory die vertically to dramatically boost bandwidth — is produced mainly by SK Hynix and Samsung, and supply is already tight.
This contrasts with recent public statements from Nvidia executives. Hardware engineering SVP Andrew Bell told media in mid-July: "We've planned ahead on memory — supply won't be a near-term constraint… pricing pressure is probably the bigger challenge."
This reflects a deeper reality: even with a lead in chip design, upstream memory-supply bottlenecks can force Nvidia to redesign its flagship product.
03

Less memory — what happens to performance?

Reduced memory does affect chip performance, but Nvidia may compensate. Two customers say faster network interconnects and more efficient data-storage methods can let users spread models and workloads across more Rubin Ultra chips.
In plain terms = if one chip doesn't have enough memory, link several together; with fast enough networking, total performance may not suffer much.
The trade-off is direct: customers running frontier-scale AI models would need to deploy more chips to match the same compute, and total spending may not fall.
04

How are customers reacting?

One customer noted that lower memory means per-chip cost could drop, making it a more affordable option for companies watching their budgets.
Another said the company cares more about its long-term, multi-generation hardware relationship with Nvidia than any single chip's memory spec. This means → for large buyers, locking into Nvidia's ecosystem matters more than one generation's parameters.
HBM can account for over half the component cost of an advanced AI chip, so lower-memory versions may carry a significant cost advantage.
05

When will the final spec be set?

Rubin Ultra is scheduled to ship by late next year; Nvidia has not finalized all specifications.
This means → Nvidia still has time to adjust the design based on memory supply, cost, and customer demand — current prototypes are not the final product.
Nvidia declined to comment.

Content is for reference only, not financial advice.

Nvidia Testing Plans to Cut HBM Capacity on Rubin Ultra Chips · nashnova