AI · Sep 15, 2026
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI FactoriesNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
Sep 16, 2026, 8:00 AM · NVIDIA Blog

In its MLPerf Inference v6.1 debut, NVIDIA says Vera Rubin NVL72 posts up to 3.7x the throughput of GB300 NVL72, with GB300 four-rack runs at 99% scaling efficiency.
Why it matters
NVIDIA’s blog on MLPerf Inference v6.1 claims Vera Rubin NVL72’s first preview submission delivers up to 3.7x better throughput than GB300 NVL72. A 288-GPU submission across four GB300 NVL72 racks hit 99% scaling efficiency. Software optimizations versus v6.0 added up to 1.6x in NVIDIA’s submissions, with further gains after the formal cutoff.
Inference economics turn on tokens per watt and near-linear scale-out. MLPerf is the public scoreboard labs and buyers cite—even when vendors submit their own stacks.
Treat the numbers as vendor-submitted until cross-checked, but the directional story is Rubin as the next rack.
From the desk
We’re reading MLPerf week the way operators do: throughput, scaling efficiency, software climb.
NVIDIA’s framing—system performance, efficient scaling, continuous optimization, plus “platform fungibility”—is marketing language with an operational core. If Vera Rubin NVL72 really lands near 3.7x GB300 NVL72 on the submitted workloads, procurement conversations shift. The 99% four-rack GB300 result is the quieter tell: buyers fear stranded GPUs more than they fear missing a peak TOPS slide.
Useful AI at scale needs racks that turn power into tokens without heroic utilization theater. We’re for hardware that makes serving cheaper and more predictable.
The harm path: benchmark shopping, preview submissions that don’t match GA configs, and CapEx stampedes that outrun grid interconnects. MLPerf is necessary; it isn’t sufficient.
I’m watching peer submissions on the same benchmarks and whether post-deadline software gains show up in customer clusters, not only blog charts.
Context
MLPerf Inference is run by MLCommons. NVL72 denotes NVIDIA’s liquid-cooled 72-GPU rack scale-up systems. Vera Rubin is the successor platform generation to Blackwell-class GB300 systems in NVIDIA’s naming.
Who feels it
- Hyperscalers and neoclouds
- Rubin preview numbers will feed 2027 capacity plans—validate against your model mix, not only NVIDIA’s.
- Software inference teams
- Up to 1.6x from software between v6.0 and v6.1 underscores stack work as a first-class lever.
- Competitors
- Must answer on throughput-per-rack and multi-rack efficiency, not only single-accelerator peaks.
What to watch
- Full MLCommons result tables versus NVIDIA’s selective highlights.
- GA Vera Rubin configs matching preview submission topology.
- Customer-reported tokens-per-watt after software patches land.
Companies: NVIDIA