NVIDIA Groq LPX Racks Enter Mass Production; SpaceX and Nebius Become First Vera CPU Customers

Nashnova编辑部
Published todayAbout 9 min read

Nvidia announced full production of Groq 3 LPX racks, with SpaceX and Nebius as launch customers — marking the commercial debut of low-latency inference hardware after Nvidia's $20 billion Groq asset acquisition.

01

What exactly is the Groq rack selling?

Each Groq 3 LPX rack packs 256 Groq 3 chips, purpose-built for low-latency inference — speeding up the "answer generation" step of AI models, not the training phase.
Benchmark tests show 3,400 tokens per second. OpenAI's new Ultrafast mode promises 750 tokens per second. This means → Groq opens a roughly 4.5× speed gap on raw inference throughput.
The core design embeds 500 megabytes of high-speed SRAM directly on-chip. In plain terms = data doesn't have to travel off-chip to fetch memory, so latency drops sharply.
02

How will SpaceX and Nebius use them?

Nebius will become the first AI cloud provider to deploy both Groq LPX racks and Vera Rubin GPU racks. The Groq racks ship alongside Vera CPUs and Rubin GPUs, going live later this year.
SpaceX plans to run Vera CPU racks — each holding up to 256 Vera CPUs — for long-running agentic AI workloads (AI that autonomously executes multi-step tasks).
SpaceX's SpaceXAI unit will also use Nvidia's Vera Rubin NVL72 systems to launch Starmind AI satellites. This reflects Nvidia's compute product line extending from data centers into space.
03

If it doesn't replace GPUs, what is the positioning?

Nvidia senior director Dion Harris stated plainly: "This isn't about replacing GPUs — it's about using the right processor at the right price for the right workload."
In plain terms = GPUs handle the heavy lifting (training and complex inference); Groq handles the fast lane (the latency-critical decoding phase). Two product lines, each covering a different stage.
This means → cloud providers can package low-latency capability as a premium service tier, charging more for the most speed-sensitive customers.
04

What are competitors doing?

AMD previously announced integrating its rack-scale systems with Cerebras chips, targeting the same low-latency inference segment. Cerebras recently completed its IPO.
Nvidia's Groq LPX mass production is a direct response to this competitive landscape. This reflects low-latency inference moving from proof-of-concept to a customer-acquisition race.
Groq chips are fabricated by Samsung; Nvidia GPUs by TSMC — two parallel supply chains that also diversify manufacturing concentration risk.
05

What hasn't been disclosed yet?

According to The Information, Nvidia has not revealed pricing, volume, or delivery timelines for Vera CPU or Groq LPX racks.
Some server vendors say they have received neither pricing information nor purchase orders. This means → commercialization is just starting; actual shipment scale remains unknown.
Nvidia reports earnings this Wednesday. Whether Groq LPX customer order volume and pricing surface in the earnings call is the market's next focal point.

Content is for reference only, not financial advice.