NVIDIA Vera Rubin NVL72's MLPerf Inference v6.1 Debut
TECH

NVIDIA Vera Rubin NVL72's MLPerf Inference v6.1 Debut

23+
Signals

Strategic Overview

  • 01.
    NVIDIA published its MLPerf Inference v6.1 submission results on September 16, 2026, headlined by the first MLPerf Inference preview submission of its Vera Rubin NVL72 system, posting up to 3.7x higher throughput than GB300 NVL72 on the Qwen3-VL benchmark.
  • 02.
    NVIDIA submitted Vera Rubin NVL72 preview results on two of the suite's most demanding benchmarks, reporting up to 2.5x higher token throughput than GB300 NVL72 on DeepSeek-R1 using TensorRT-LLM, and up to 3.7x on Qwen3-VL using vLLM with NVIDIA Dynamo.
  • 03.
    The round included five new processors or accelerators making their MLPerf debut, with Vera Rubin NVL72 designated a preview (not-yet-commercially-available) entrant among a record 30 submitting organizations and 486 datacenter and edge results.
  • 04.
    Vera Rubin's gains stem from enhanced Tensor Cores and Transformer Engine accelerating both prefill and decode stages, plus NVFP4 precision, which cuts the memory footprint of model weights, attention, and KV cache.

Deep Analysis

Why Vera Rubin Wins: NVFP4 Precision Meets Purpose-Built Silicon

Vera Rubin NVL72's throughput lead over GB300 NVL72 traces to a combination of hardware and software changes rather than any single trick. NVIDIA describes the system's enhanced Tensor Cores and Transformer Engine as accelerating both the prefill and decode stages of inference, while NVFP4 precision shrinks the memory footprint of model weights, attention, and KV cache - raising throughput with what the company calls minimal loss of output quality [1]. Those software gains ride on a rack that packs 72 Rubin GPUs and 36 Vera CPUs with 20.7 TB of HBM4 GPU memory, 1,400 TB/s of GPU memory bandwidth, and 216 TB/s of NVLink bandwidth, rated at 3,600 PFLOPS of NVFP4 inference compute [2]. The DeepSeek-R1 result (2.5x) ran on TensorRT-LLM across offline, server, and interactive scenarios, while the larger Qwen3-VL gain (3.7x) used vLLM paired with NVIDIA Dynamo - a reminder that the advertised leap is really two separate stacks tuned for different workload shapes, not one universal multiplier.

The Preview-Category Playbook: Why Benchmark Hardware That Isn't Shipping Yet

MLPerf's preview category exists for platforms expected to reach commercial availability by the next submission round, and NVIDIA used it to put Vera Rubin NVL72 numbers on the record while GB300 NVL72 remains the only Rubin-era product customers can actually buy today [3]. That is a deliberate sequencing: publish third-party-verified proof of the next platform's performance ceiling before general availability, so cloud buyers can plan capital allocation around it rather than around marketing slides alone. Nebius is the clearest evidence the strategy is landing with real customers - the cloud provider ran its own Vera Rubin NVL72 preview submission and has committed to offering the platform commercially in the US and Europe from the second half of 2026 [4]. It is also a contrast play: NVIDIA measures Vera Rubin's preview numbers directly against GB300 NVL72's own MLPerf v6.0 record - a 288-GPU, four-rack run that pushed 2.5 million tokens per second on DeepSeek-R1 and was, at the time, the largest submission scale in the benchmark's five-year history [5]. Vera Rubin's preview debut is framed to make that record look like the previous generation almost immediately.

AMD and Intel Narrow the Gap Even as NVIDIA Leaps Ahead

NVIDIA's preview numbers are the headline, but the same v6.1 round quietly shows how much closer AMD and Intel have gotten to NVIDIA's shipping hardware. AMD posted its first MLPerf results for Instinct MI350P and Ryzen AI Max+ 395, and partnered with Crusoe to run 512 Instinct MI355X GPUs across 64 nodes on standard RoCE Ethernet - the largest system ever submitted to MLPerf Inference, delivering 5.75 million tokens per second offline and 5.39 million tokens per second in the server scenario on gpt-oss-120b [6]. Intel, meanwhile, made its own MLPerf debut with Arc Pro B70 GPUs, posting generational gains across Llama 3.1 8B, Llama 2 70B, gpt-oss-120B, Whisper, and the new End-to-end RAG benchmark. The round drew a record 30 submitting organizations and 486 datacenter and edge results in total [7], and MLCommons added the End-to-end RAG test to this cycle specifically because deployment patterns have moved past single-turn query-answering, which is exactly the terrain Vera Rubin NVL72 is architected for [6]. NVIDIA is still ahead, but the field it is racing is visibly closing distance at the same moment it previews its next platform.

Reading the Fine Print: When a Triple-Digit Multiple Isn't the Real Number

The most attention-grabbing figure from this round is not a MLPerf metric at all: Vera Rubin NVL72 posted a 30x performance gain over GB300 NVL72 on the SemiAnalysis AgentX benchmark, a third-party test designed specifically to measure agentic workloads [8]. Numbers at that scale invite exactly the kind of skepticism that surfaced in investor-focused discussion of the results, where a thread summarizing a related SemiAnalysis performance-per-dollar analysis cited far larger multiples - tens of times more tokens per dollar of total cost of ownership than GB300 at high-interactivity service levels. A commenter in that same thread pushed back that those headline multiples 'only apply at high interactivity where the base is very low,' meaning the eye-catching ratios compress dramatically once you look at typical, moderate-throughput deployment rather than the narrow edge case that produces the biggest number. That tension - real, verified architectural gains on one side, and best-case marketing multiples on the other - is the honest way to read every number in this launch: the 2.5x and 3.7x MLPerf figures are peer-reviewed and reproducible, but the more dramatic double- and triple-digit multiples circulating around Vera Rubin depend heavily on which slice of the workload curve gets measured.

Historical Context

2026-01
Unveiled the Rubin platform, including Vera Rubin NVL72, at CES 2026 as the successor to Blackwell.
2026-03
Set MLPerf Inference v6.0 records with a 288-GPU, four-rack submission pushing 2.5 million tokens per second on DeepSeek-R1, the largest submission scale in the benchmark's five-year history at that time.
2026-09-16
Posted its first preview-category MLPerf Inference results in the v6.1 round, alongside first-time submissions from AMD Instinct MI350P, Intel Arc Pro B70, and AMD Ryzen AI Max+ 395.

Power Map

Key Players
Subject

NVIDIA Vera Rubin NVL72's MLPerf Inference v6.1 Debut

NV

NVIDIA

Designed the Vera Rubin NVL72 and GB300 NVL72 hardware plus the TensorRT-LLM, Dynamo, and NVFP4 software stack being benchmarked, and submitted both platforms' results.

ML

MLCommons

Independent body that runs the MLPerf Inference benchmark, verified this round's submissions across 30 organizations, and added new tests reflecting shifting AI deployment patterns.

AM

AMD

Submitted first MLPerf results for Instinct MI350P and Ryzen AI Max+ 395, and partnered with Crusoe on the round's largest GPU cluster submission (512 Instinct MI355X GPUs across 64 nodes).

IN

Intel

Submitted first MLPerf results for Arc Pro B70 GPUs, posting generational gains across Llama 3.1 8B, Llama 2 70B, gpt-oss-120B, Whisper, and End-to-end RAG benchmarks.

NE

Nebius

Cloud partner that submitted its own Vera Rubin NVL72 preview results and plans to offer the platform commercially in the US and Europe starting in the second half of 2026.

Fact Check

8 cited
  1. [1] NVIDIA Blog: Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
  2. [2] NVIDIA: Vera Rubin NVL72 Data Center Platform
  3. [3] MLCommons: MLPerf Inference v6.1 Results
  4. [4] Nebius Newsroom: Nebius to Offer NVIDIA Vera Rubin NVL72 in US and Europe from H2 2026
  5. [5] NVIDIA Blog: MLPerf Inference Blackwell Ultra (v6.0)
  6. [6] StorageReview: MLPerf Inference v6.1 - 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin's First Peer-Reviewed Numbers
  7. [7] GlobeNewswire: MLCommons Sets Participation Record With New MLPerf Inference v6.1 Benchmark Results
  8. [8] Unite.AI: NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results

Source Articles

Top 3

THE SIGNAL.

Analysts

Framed the new End-to-end RAG test as reflecting how AI deployment scenarios have shifted beyond simple LLM query-answering: 'We added the End-to-end RAG test because it's clear that query-answering has evolved beyond simply an LLM trained on a corpus.' He added that the benchmark suite's goal is to track what the field actually cares about: 'We are working hard to ensure that the MLPerf Inference benchmark continues to reflect the scenarios that the AI community values most.'

Miro Hodak
MLPerf Inference working group co-chair, MLCommons

Positioned purpose-built infrastructure as a requirement for the agentic AI era: 'Leading in the era of agentic AI requires infrastructure that is purpose-built for scale, performance, reliability and cost efficiency.'

Dave Salvator
Director of accelerated computing products, NVIDIA
The Crowd

Five first-place results in MLPerf® Inference v6.1. 603,023 tokens/sec on DeepSeek R1 at 72 GPUs and preview-category results on the Nebius @nvidia Vera Rubin NVL72. We were one of only two submitters with results on that hardware. Full results: nebius.com/blog/posts/mlperf-inference-v6-1-results #MLPerf

@@nebiusai158

Nvidia Vera Rubin NVL72 debuts in MLPerf Inference In its first MLPerf Inference v6.1 preview, Nvidia said Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL, and up to 2.5x on DeepSeek-R1

@@onlycapex0

NVIDIA's Vera Rubin NVL72 just debuted in MLPerf Inference v6.1: up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL, and up to 2.5x on DeepSeek-R1. Preview results posted today. https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/

@@ProLogicaAI0

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

@u/bl079725
Broadcast
NVIDIA Unveils Vera Rubin: The World's Most Powerful AI Supercomputer | CES 2026 NVIDIA | AI14

NVIDIA Unveils Vera Rubin: The World's Most Powerful AI Supercomputer | CES 2026 NVIDIA | AI14

NVIDIA Vera Rubin NVL72 production racks are here.

NVIDIA Vera Rubin NVL72 production racks are here.

2026 Best Choice of the Year: NVIDIA Vera Rubin NVL72 - The Peak of AI Supercomputing

2026 Best Choice of the Year: NVIDIA Vera Rubin NVL72 - The Peak of AI Supercomputing