Vera Rubin’s real-world performance revealed for the first time, showing a significant boost! Nvidia “crashes” AMD’s conference ahead of time
Nvidia made intensive announcements on the eve of AMD's annual conference: the Vera CPU nearly doubles the performance of x86 chips and compresses latency by six times, while real-world tests show the Vera Rubin NVL72 achieves ten times the energy efficiency of the previous generation. OpenAI, Anthropic, and SpaceX have already received the first batch of chips. This carefully timed data blitz is not only a direct hit against AMD, but also serves as Nvidia's declaration of intent to aggressively enter the CPU market and aim for a $200 billion pie.
On the eve of rival AMD's annual product launch event, Nvidia landed a heavy blow by intensively disclosing real-world performance data of its next-generation Vera Rubin platform and officially unveiling the full specifications of its self-developed CPU chip, Vera.
On Tuesday, Nvidia revealed that Vera CPU delivers nearly twice the performance of x86 chips on agent-based AI tasks with latency improved by as much as six times. Meanwhile, early production tests by cloud partner CoreWeave showed that, when running the DeepSeek R1 model, the Vera Rubin NVL72 platform produced 10 times the token throughput per megawatt of computing power compared to the previous-generation Blackwell-based GB200 NVL72 system.
The timing of this data release is deeply significant. AMD will hold its annual product launch event, "Advancing AI," in San Francisco this Thursday, and Nvidia's concentrated release of performance data has widely been interpreted as a deliberate market maneuver. For cloud vendors and enterprise customers evaluating next-generation AI infrastructure investments, this batch of data will directly influence their procurement decisions.
At the same time, Nvidia also officially announced that the Vera CPU completed its first batch of deliveries in June, with customers including OpenAI, Anthropic, and SpaceX. This marks Nvidia's formal move into the CPU market, making a direct assault on the traditional strongholds of AMD and Intel.
Vera CPU Unveiled: Nvidia Officially Enters the Server CPU Market
On Tuesday, Nvidia announced the full specifications, benchmark results, and architectural details of its data center CPU product, Vera—critical data needed by potential customers for comprehensive evaluation. The company stated that the Vera chip completed delivery to clients, including OpenAI, Anthropic, and SpaceX, in June.
Vera is Nvidia's first server CPU to be independently designed from the core level, featuring a custom microarchitecture codenamed "Olympus core" rather than relying on Arm's off-the-shelf designs.
Hannah Coutand, Nvidia's Product Marketing Lead for Vera, introduced that the chip's design focuses on single-core speed, high memory bandwidth, and low latency, with the goal of "enabling agents to return to the GPU as quickly as possible and keeping the GPU highly utilized at all times." In terms of hardware specifications, Vera has a TDP ranging from 250 watts to 450 watts and supports up to 1.5TB of low-power memory per chip.
In the latest benchmarks, Nvidia demonstrated that Vera delivers 1.9x the performance of x86 chips on agent-based AI tasks and reduces latency by a factor of six. Nvidia also stated that Vera surpassed AMD’s flagship EPYC Turin CPU by nearly 100% in certain industry-standard benchmarks. Previously, Nvidia claimed Vera outperformed x86 chips by 50% in overall AI agent performance.
The launch of Vera marks Nvidia's latest move in its vertical integration strategy. Wolfe Research estimates the average selling price of a Vera chip to be around $5,000, with shipment volume projected to reach about 1.3 million units this year. Ian Buck, Nvidia’s VP of Hyperscale Computing, said agent AI has made CPUs “even more indispensable” and predicted that the total server CPU market could ultimately reach $200 billion.
Vera Rubin In Real-World Testing: 10x Energy Efficiency Leap
In early June, CoreWeave completed the industry’s first deployment and validation of the Vera Rubin NVL72 platform—including full-stack confirmation of power, cooling, networking, and computing. The disclosed data represents the first public, real-world results for Vera Rubin NVL72 silicon performance.
The tests, based on the DeepSeek R1 model, showed that at identical interaction response targets, Vera Rubin NVL72 generated 10 times as many tokens per megawatt per second as GB200 NVL72. CoreWeave also noted that optimizations made for Vera Rubin can be retroactively applied to GB200 NVL72 systems, increasing the latter’s throughput per megawatt more than fourfold in just three months. Nvidia stated these achievements have been validated through over 250,000 unique configurations and more than 1.4 million GPU hours of testing.
Architecturally, the Vera Rubin NVL72 rack integrates 72 Rubin GPUs and 36 Vera CPUs, interconnected via 260 TB/s full-mesh NVLink 6 and supporting native NVFP4 precision. CoreWeave said this efficiency boost allows customers to handle more inference traffic within the same power budget or to run the same workload with less energy, reducing the cost per token.
CoreWeave also revealed concrete application scenarios: a global cybersecurity firm expects to run threat detection inference at 10 times the per-watt performance; an autonomous programming agent company plans to scale agent workloads at dramatically lower token costs; and an AI-native search engine expects to expand real-time search services to more users without breaching response time limits.
Infrastructure Synergistic Optimization: Hardware-Software Integration Unleashes More Efficiency
Nvidia emphasized that these performance improvements aren’t solely the result of the chip itself, but come from the comprehensive synergy of hardware and software co-design. Through dynamic optimization of the entire infrastructure and energy stack, Nvidia stated it can deploy 40% more GPUs within the same power envelope.
On the cooling side, Nvidia has adopted a 45°C closed-loop liquid cooling system, saving around 4 million gallons of water per megawatt annually compared to standard cooling methods. On the networking side, the sixth-generation NVLink 6 interconnect delivers 2.3x higher throughput for large language model inference decoding than Ethernet; the Spectrum-X platform achieves 1.6x faster remote direct memory access bandwidth, cuts switch count by 1.7 times, boosts optical power efficiency fivefold, and improves reliability tenfold.
Nvidia’s latest generation Spectrum-6 platform has already begun shipping to AI factory clients such as CoreWeave, Microsoft, Nebius B.V., SpaceXAI Corp., and Tesla. Currently, Nvidia is ramping up shipments to customers and partners including Google Cloud, Microsoft Azure, Meta, Oracle Cloud Infrastructure, Dell Technologies, OpenAI, and CoreWeave.
Nvidia Remains a Challenger in the CPU Space
Despite Vera’s impressive performance data, Nvidia still faces significant market share challenges in the CPU arena. Gartner analyst Kevin Knox pointed out that AMD is currently the leading competitor in the enterprise AI server CPU field. Reports indicate that Intel holds about 66.8% of the server CPU market, AMD about 33%, with AMD steadily gaining ground and establishing deep partnerships with hyperscale cloud providers.
Driven by CPU demand from the rise of agent-based AI, AMD’s and Intel’s stock prices have surged 128% and 149% respectively year-to-date, far outpacing Nvidia’s roughly 8% gain in the same period, making them the standout chip stocks for 2026.
For now, Nvidia's list of major cloud service partners for Vera includes only Oracle, not other mainstream cloud vendors. Hannah Coutand admitted that Vera is still in the "early adoption phase," but OpenAI plans to begin large-scale deployment of Vera chips as early as this quarter.
Karl Freund, founder of Cambrian AI Research, summarized Nvidia’s strategic intent: "Their aim is to free customers from dependency on Intel or AMD CPUs, and they are determined to capture this revenue. They have chosen to focus on building a unique CPU that no one else in the market is currently offering."
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like
HYPE corrects to $77 as traders watch $70-$75 support after $90 rejection
BCA: Liberation Day 2.0 - Tariff Policy and Outlook for the 2026 Midterm Elections

Ripple partners with SettleMint to streamline bank asset tokenization
RippleX explores XRP as institutional collateral with XLS-65 and XLS-66
