Bitget App
Trade smarter
Buy cryptoMarketsTradeFuturesEarnAISquareMore
As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months

华尔街见闻华尔街见闻2026/07/22 08:27
Show original
By:华尔街见闻

NVIDIA is extending the lifecycle of the Blackwell platform through continuous software stack optimization, while also preparing for the large-scale deployment of its next-generation Rubin architecture.

Against the backdrop of accelerated deployment of the new Rubin platform, NVIDIA revealed that it has quadrupled the per-megawatt throughput (TPS/MW) of the GB200 NVL72 running the DeepSeek R1 0528 model in just three months. This achievement is the result of 38 major optimizations completed by NVIDIA within four months, backed by over 250,000 simulated configuration tests and a cumulative total of 1.4 million GPU hours of optimization validation.

At the same time, the GB300 "Blackwell Ultra" platform continues to break records in AI training benchmarks. At a scale of 256 cards, the GB300 NVL72 set a new all-time high in the DeepSeek-V3 671B pre-training task with a throughput of 1,648 TFLOPs per GPU—about three times the 606 TFLOPs achieved by the GB200, and this metric has already increased 1.5 times over the past six months.

Significant Energy Efficiency Leap for GB200, Optimizations Cover Over 90% of Models

NVIDIA stated that a series of optimizations for the GB200 NVL72 platform have resulted in a fourfold increase in per-megawatt throughput when running DeepSeek R1 0528 (1K input/1K output configuration) over three months.

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months image 0

This improvement is enabled by 38 major optimization iterations completed by NVIDIA within four months. These optimizations were selected through more than 250,000 simulated configuration tests and validated with a cumulative total of 1.4 million GPU hours. NVIDIA specifically noted that over 90% of these optimizations are reusable across models in its AI product portfolio, rather than being tailored for a single task.

This means that existing data center customers can achieve significant energy efficiency gains through software upgrades alone, without hardware replacement—a direct economic benefit for hyperscale cloud providers and enterprise customers who are highly sensitive to computing costs.

GB300 Breaks Training Records, 1.5x Performance Increase Over Six Months

On the AI training side, the GB300 NVL72 also demonstrates robust performance momentum. At a 256-card scale, using the Megatron Core framework for pre-training the DeepSeek-V3 671B model, the GB300 NVL72 achieved a throughput of 1,648 TFLOPs per GPU, roughly three times the previous GB200’s 606 TFLOPs, setting a new global record for this task.

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months image 1

Notably, this figure is not a static hardware limit. GB300 NVL72’s performance under the Megatron Core framework grew from 1,088 TFLOPs/GPU in November 2025 to 1,648 TFLOPs/GPU in June 2026, an approximate 1.5x increase within six months.

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months image 2

In terms of collaborative optimization with mainstream AI frameworks, NVIDIA’s deep cooperation with the PyTorch and JAX communities has also brought significant gains. On TorchTitan (PyTorch native training stack), GB300 NVL72’s training performance on DeepSeek-V3 671B improved sixfold compared to the unoptimized baseline, rising from 199 TFLOPs/GPU to 1,197 TFLOPs/GPU. The improvement was even greater under the JAX framework; as of July 2026, per-GPU throughput reached 4,082 Tokens/s, corresponding to 1,025 TFLOPs/GPU—approximately ten times the 418 Tokens/s achieved in January 2026.

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months image 3

Scaling Efficiency Near Theoretical Limits, 800 Gb/s Networking Is Key

Scaling efficiency in large-scale training scenarios has always been a key metric for evaluating the practicality of AI infrastructure. NVIDIA disclosed that within the scaling range of 256 to 1,024 cards, the GB300 NVL72 maintained scaling efficiency close to the theoretical limit across three major frameworks: Megatron Core at 98.5%, and both TorchTitan and JAX at 97%.

NVIDIA attributes this scaling performance to the 800 Gb/s Scale-Out networking chip built into the NVL72 rack. High-speed interconnects directly affect communication overhead in multi-server training, thereby impacting overall scaling efficiency. This network capability is regarded as the core infrastructure component supporting such high cross-framework efficiency.

As Rubin is highly anticipated, NVIDIA continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months image 4

Rubin Platform Accelerates Rollout, Blackwell Optimizations Ongoing

These optimization advances are announced as NVIDIA’s next-generation Vera Rubin platform enters global deployment. Reportedly, the Vera Rubin NVL72 delivers around a tenfold increase in Token throughput compared to Blackwell: GB200 NVL72 reaches about 80,000 Tokens/s, while Vera Rubin NVL72 can achieve up to 800,000 Tokens/s at the same 150MW power level.

However, as the Blackwell platform is already deployed in countless data centers worldwide, NVIDIA’s strategy mirrors that of the Hopper generation—continuing to unlock the potential of deployed hardware through software optimization while advancing the new platform. This approach not only extends the ROI cycle of current customers’ hardware investments, but also reinforces NVIDIA’s platform stickiness in the AI infrastructure ecosystem. For the market, beneath NVIDIA's dual-platform layout with Blackwell and Rubin, it is gradually building a software moat that competitors will find hard to surpass in the short term.

0
0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.

Understand the market, then trade.
Bitget offers one-stop trading for cryptocurrencies, stocks, and gold.
Trade now!