memujo
AI6 min read

Vera Rubin NVL72's 3.7x Qwen3-VL MLPerf Lead

Vera Rubin NVL72 posts its first MLPerf Inference v6.1 results, up to 3.7x Qwen3-VL throughput over GB300 and 99% four-rack scaling. Here is the per-GPU math.

By Alice

In this article
  1. 01The facts, at list price
  2. 02The math, worked out
  3. 03How the gain is actually made
  4. 04The number that matters for operators
  5. 05What the pitch leaves out
  6. 06The field is moving
  7. 07Our Read

NVIDIA's Vera Rubin NVL72 just posted its first verified MLPerf Inference numbers, and the headline is a 3.7x throughput gain over the outgoing GB300 NVL72. That number is real, but it is the peak, not the average. Pull the actual tokens per second out of the v6.1 results and the generational step is closer to 2x per GPU, which changes what this hardware is actually worth to the person buying a rack.

The facts, at list price

On September 16, 2026, MLCommons published MLPerf Inference v6.1, a round that set a participation record with 30 submitting organizations and 486 datacenter and edge results, according to StorageReview. Two new benchmarks joined the suite: an End-to-End RAG pipeline for the datacenter and an Edge Agentic Inference test for single-user devices. For the first time, the results include verified numbers for NVIDIA's Vera Rubin NVL72, AMD's Instinct MI350P, and Intel's Arc Pro B70.

NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the suite, DeepSeek-R1 and Qwen3-VL, and reported the headline figures in its own results blog. On Qwen3-VL, the company reported up to 3.7x higher throughput than GB300 NVL72 across the offline, server, and interactive scenarios, using vLLM with the NVIDIA Dynamo framework. On DeepSeek-R1, using TensorRT-LLM, it reported throughput up to 2.5x higher. Both are "up to" figures, which is the first caveat to hold onto.

The math, worked out

The raw numbers let us check that headline. In the DeepSeek-R1 server scenario, a full 72-GPU Vera Rubin NVL72 rack posted 1,175,890 tokens per second, against 603,023 tokens per second for a 72-GPU GB300 NVL72. That is a 1.95x ratio, not 3.7x. In the offline scenario the gap narrows to 1.72x, with Vera Rubin at 1,183,326 and GB300 at 689,961 tokens per second.

The 3.7x belongs to Qwen3-VL, a vision-language model, and it is the best case across three scenarios. The 2.5x is the DeepSeek-R1 best case. Neither is a single clean number, and the per-model spread is the whole point: the generational gain is not uniform across workloads.

The cleanest like-for-like comparison is per GPU, and that is exactly how Nebius, one of only two submitters with Vera Rubin hardware, framed its own results. In its v6.1 writeup, Nebius ran DeepSeek-R1 on a 36-GPU Vera Rubin configuration and posted 16,427 tokens per second per GPU in the server scenario, against 8,375 tokens per second per GPU for a 72-GPU GB300 NVL72 (603,023 total). That is a 1.96x per-GPU gain, which matches the 72-versus-72 math almost exactly. The honest, workload-averaged generational step on a frontier reasoning model is roughly 2x per GPU.

How the gain is actually made

NVIDIA attributes the results to full-stack codesign, and the specific levers are worth naming because they explain why the per-GPU number moved at all. The NVL72 scale-up domain runs on sixth-generation NVLink and NVLink Switch, which NVIDIA says delivers 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet. That interconnect is what makes the two techniques that dominate these workloads work at rack scale.

The first technique is disaggregated serving, which separates the prefill and decode stages of inference onto different hardware so each can be tuned independently. The second is large-scale expert parallelism across the mixture-of-experts layers that power both DeepSeek-R1 and Qwen3-VL. On top of that, NVFP4 precision reduces the memory footprint of model weights, attention, and the KV cache, which raises throughput with what NVIDIA describes as minimal loss of output quality. None of these is a chip-level feature you can see in a spec sheet, and together they are why a 2x per-GPU jump is achievable without a 2x jump in the transistor count.

The number that matters for operators

The other figure NVIDIA leans on is scaling efficiency, and it is the one that matters most to a datacenter operator. NVIDIA's DeepSeek-R1 submission scaled from a single 72-GPU GB300 NVL72 rack to four racks, 288 GPUs, at 99% scaling efficiency in the offline scenario. Four times the hardware produced 3.96x the throughput, which is essentially linear. Nebius independently measured the same property on gpt-oss 120B: moving from 8 to 72 GB300 GPUs, a 9x increase, delivered 8.8x server throughput and 9.0x offline throughput. Adding GPUs still buys you nearly proportional throughput, which is the property that keeps a large cluster from wasting money.

What the pitch leaves out

Here is where the story gets more complicated. In absolute aggregate throughput, a single Vera Rubin NVL72 rack does not top this round. AMD's 512-GPU MI355X cluster posted 2,901,950 tokens per second on DeepSeek-R1 offline and 2,405,310 in the server scenario, more than double what one Vera Rubin rack produces. The crown for raw aggregate throughput still goes to whoever stacks the most GPUs, not to the newest chip. What Vera Rubin changes is the efficiency of each GPU, not the ceiling of a single rack.

There is also a caveat on the pricing side. NVIDIA frames the results as lowering cost per token, and the direction is right. If a Vera Rubin rack produces roughly 1.95x the tokens of a GB300 rack, then at equal rack cost the cost per token falls to about 51% of the GB300 level, or roughly half. That is an assumption, though. I am assuming the two racks cost the same to buy and run, and NVIDIA has not published a Vera Rubin NVL72 price, so the actual cost-per-token delta depends on a number that is not in these results. The 3.7x Qwen3-VL figure would imply an even steeper drop on that model, but again it is the peak, not the blended number.

Finally, the most aggressive number in the whole release is the one that is not MLPerf-verified. NVIDIA says Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 on the SemiAnalysis AgentX benchmark, which measures agentic workloads, in preview testing. That is a large claim, it is not in the MLCommons results, and it should be read as a preview figure from the vendor, not a verified score. The company's point is that agentic workloads, where a model reasons, plans, and acts across many steps, are reshaping how inference is measured, and that the upcoming MLPerf Endpoints benchmark will bring standardized measurement to them.

The field is moving

For perspective on the pace, MLCommons says the best per-accelerator DeepSeek-R1 server result in v6.1 is 5.7x better than the best result in v5.1 a year ago, and the best vision-language result improved 2.99x in the six months since v6.0. A 2x per-GPU generational jump on top of a 5.7x year-over-year field improvement is why the absolute numbers keep climbing even when the per-generation ratio looks modest. Software is a big part of that: NVIDIA reports its GB300 NVL72 Qwen3-VL performance improved up to 1.6x over v6.0 through lower KV cache precision, additional kernel fusion, and better kernels, with further post-submission gains on GPT-OSS-120B and DLRMv3 that are not yet verified by MLCommons. The round also drew 19 partners, eight of them on multi-node Blackwell NVL72 systems, which is a signal that the ecosystem is already shipping on this stack.

Our Read

Vera Rubin NVL72 is a genuine, independently verified roughly 2x per-GPU generational leap, and the 3.7x headline is a real but model-specific peak that should not be read as the average. For a buyer the practical takeaway is cost per token, and the verified math says roughly a halving relative to a GB300 rack, with the caveat that the Vera Rubin price is not published. The strategic takeaway is different from what the press release implies. This is not a monopoly moment. In aggregate throughput, a 512-GPU AMD cluster still outproduces a single Vera Rubin rack, and the field is moving fast enough that a 2x per-GPU gain is the new normal, not the ceiling. The step change is real, but it is a step, not a wall.

See also: GPT-6 Astra's Hidden Edge Is Output Tokens, Not Accuracy

  • #vera-rubin
  • #mlperf
  • #nvidia
  • #ai-inference
  • #benchmarks

Sources

Share this story