NVIDIA GB200 Energy Efficiency Quadruples in Three Months as GB300 Training Records Continue to Be Broken

Claire Weston
Published todayAbout 11 min read

Nvidia quadrupled the GB200 NVL72's throughput per megawatt in just three months, while the GB300 set a new global training record — deployed customers can capture the gains through software upgrades alone, with no hardware swap.

01

How did GB200 efficiency jump 4× in three months?

Running the DeepSeek R1 0528 model, the GB200 NVL72's tokens-per-second per megawatt (TPS/MW) rose 4× in three months.
This means → customers who already own Blackwell racks need not replace any hardware — a software update alone quadruples the work each megawatt delivers.
Behind the gain: 38 major optimizations completed in four months, filtered from over 250,000 simulated configurations and validated with 1.4 million GPU-hours of real testing.
In plain terms = Nvidia industrialized the tuning process — mass simulation narrows the field, then real hardware confirms each winner.
02

Do these optimizations only work for one model?

Nvidia stated explicitly that over 90% of the gains transfer across models, covering its broader AI product portfolio.
This means → this is not benchmark-specific tuning for DeepSeek — it is a foundational, portable improvement. Switch models, and most of the benefit stays.
This reflects a strategic shift from "top the leaderboard on one model" to platform-wide energy efficiency, aimed at locking in long-term customer stickiness.
03

How much faster is GB300 at training?

At 256-GPU scale, the GB300 NVL72 hit 1,648 TFLOPs per GPU pre-training DeepSeek-V3 671B — roughly 3× the GB200's 606 TFLOPs — a new global record.
The number is still climbing: from 1,088 TFLOPs/GPU in November 2025 to 1,648 TFLOPs/GPU in June 2026, a ~1.5× gain in six months.
In plain terms = the same GPU now crunches one-and-a-half times what it could six months ago — purely through software and framework improvements.
04

How much does performance vary across training frameworks?

Megatron Core — Nvidia's own stack — leads at 1,648 TFLOPs/GPU.
TorchTitan — PyTorch's native training stack — jumped from an unoptimized baseline of 199 to 1,197 TFLOPs/GPU, a 6× gain.
JAX saw the steepest climb: per-GPU throughput rose from 418 Tokens/s in January 2026 to 4,082 Tokens/s in July — roughly 10×.
This means → whichever mainstream framework a customer runs, Nvidia is optimizing in parallel — no "our stack only" lock-in.
05

Does efficiency collapse when you scale past a thousand GPUs?

Scaling from 256 to 1,024 GPUs, all three frameworks held near-theoretical efficiency: Megatron Core at 98.5%, TorchTitan and JAX each at 97%.
Nvidia credits the NVL72 rack's built-in 800 Gb/s Scale-Out network chip — a dedicated chip for high-speed data transfer between racks.
In plain terms = the more GPUs you add, the worse the "traffic jam" between them. Nvidia's built-in high-speed network keeps the road wide enough that congestion wastes less than 3% of total compute.
06

With next-gen Vera Rubin arriving, will Blackwell be abandoned?

The Vera Rubin NVL72 delivers roughly 10× Blackwell's token throughput: the GB200 manages about 80,000 Tokens/s, while Vera Rubin reaches 800,000 Tokens/s at the same 150 MW power draw.
Yet Nvidia's playbook mirrors the Hopper era: push the new platform forward while continuing to optimize the deployed one.
This means → Blackwell buyers will not be left behind — Nvidia extends deployed hardware's useful life through ongoing software upgrades, and that is the core logic of its platform stickiness.
This reflects that what Nvidia truly sells is not just chips but a subscription-like compute service: hardware plus continuous optimization.

Content is for reference only, not financial advice.

NVIDIA GB200 Energy Efficiency Quadruples in Three Months as GB300 Training Records Continue to Be Broken · nashnova