NVIDIA Revives Rubin CPX Inference Chip, Targeting Long-Context Prefill Costs

nashnova research
今天发布阅读约 9 分钟

Nvidia has revived Rubin CPX — a chip project the market assumed was shelved — redesigning it as a dedicated accelerator for AI inference prefill. Analyst Ming-Chi Kuo expects mass production in Q1 2027, signaling Nvidia now treats inference cost as a standalone hardware problem.

01

Why build a chip just for "reading the prompt"?

As large-model context windows grow, over 50% of inference compute now goes to prefill — the stage where the model ingests all input text at once and generates a KV cache (a temporary memory of what it just read).
This means → prefill is no longer a warm-up step; it is the single largest cost center in inference.
In plain terms = when you talk to an AI, it spends more money "reading your question" than "writing the answer" — Nvidia decided to build a dedicated chip to cut the cost of reading.
02

How do the new CPX specs differ from the old plan?

Compute is now close to a standard Rubin GPU, with peak power also at 2,300 W. Memory shifts from 128 GB GDDR7 to 168 GB HBM4 — a clear upgrade, though still below the standard Rubin's 288 GB HBM4.
An 8-GPU compute tray totals roughly 1.34 TB of HBM4, enough to cover most long-context prefill and KV cache workloads.
This reflects a deliberate trade-off: CPX is not chasing raw general compute but optimizing the cost-per-token sweet spot between memory capacity and prefill efficiency.
03

Why move from a shared rack to a standalone one?

The earlier design had CPX sharing a rack with Rubin GPUs. The new version uses a standalone MGX ETL rack, letting customers scale CPX capacity independently.
Deployment options: 64, 128, 192, or 256 CPX GPUs. Each 64-GPU rack module holds 8 compute trays (8 GPUs each) plus one switch tray.
In plain terms = customers no longer have to buy an entire full-size Rubin rack to get CPX — they can add modules like building blocks, which sharply lowers the procurement threshold.
04

Why is interconnect bandwidth deliberately "downgraded"?

Inside a tray, 8 CPX GPUs connect via NVLink at 1–1.5 TB/s per GPU — roughly 40% of a standard Rubin GPU's 3.6 TB/s.
Tray-to-tray links use Spectrum-6 Ethernet over copper; cross-rack connections use OSFP optical fiber — a tiered design that steps down bandwidth at each level.
This means → CPX is not chasing full-path high bandwidth. Instead, it sizes interconnect to what prefill actually needs, pushing overall cost-performance further down.
05

How do CPX and standard Rubin GPUs work together?

Nvidia recommends a 1:1 deployment ratio: CPX handles prefill and generates the KV cache, then ships it via Ethernet RDMA to a Vera Rubin NVL72 GPU, which handles decoding (generating the reply).
In plain terms = CPX is the "reading specialist," Rubin GPU is the "writing specialist." Splitting the job across two chips costs less than making one chip do everything.
Kuo calls it "the best cost-performance option for long-context prefill" — but whether it delivers on that promise hinges on real-world results after mass production begins in 2027.

市场有风险,内容仅供研究参考,不构成投资建议。