AI · Sep 16, 2026
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 DebutHow NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
Sep 15, 2026, 9:55 AM · NVIDIA Developer

NVIDIA details Groq 3 LPX deterministic execution on Vera Rubin—cycle-exact scheduling across 256 LPUs, plus PEP/CPS tricks cutting voltage droop over 60% for interactive inference efficiency.
Why it matters
An NVIDIA Developer Blog by Weidman, Vayalapra, and Raghavan argues Vera Rubin prioritizes performance per watt. Factory-level DSX MaxLPS can shift power between racks to recover stranded capacity—up to 40% more GPUs and 35% higher token throughput in the same power envelope, NVIDIA says. Rack capacitors and Intelligent Power Smoothing absorb spikes. Groq 3 LPX’s deterministic model schedules compute and data movement across 256 LPU chips cycle-exactly; Preemptive Power and Clock Period Synthesis cut voltage droop by over 60% and shrink guardband for low-double-digit percent power reduction.
Interactive inference is spiky. Determinism turns spikes into something power engineers can plan around.
From the desk
We’re interested in the determinism thesis more than any single percentage.
If execution is cycle-exact, current demand becomes predictable; PEP/CPS can lower guardbands safely. That’s real silicon craft. Paired with DSX MaxLPS shifting power across racks, NVIDIA is selling a whole-factory control loop—not only a faster chip.
Vendor percentages (40% more GPUs, 35% more tokens, >60% droop cut) need independent racks to confirm. Useful AI serving benefits if interactivity gets cheaper per watt. The harm of overclaim: utilities and CFOs greenlight denser sites on slideshow physics.
I’m watching customer case studies that cite MaxLPS GPU-count gains under contractual power caps.
Context
Groq 3 LPX refers to NVIDIA’s integration of Groq-originated LPU deterministic execution ideas on the Vera Rubin platform per this technical blog. DSX MaxLPS is factory power-scheduling software.
Who feels it
- Power and facilities engineers
- Deterministic current curves could change how AI racks are provisioned versus peak-guessing.
- Inference product teams
- High-interactivity small-batch serving is the beneficiary if guardbands fall.
- Competitors
- Must answer with their own power-smoothing and deterministic-scheduling stories.
What to watch
- Independent measurement of droop reduction and guardband savings.
- Production MaxLPS deployments claiming +40% GPU fit under fixed MW.
- Latency/quality tradeoffs when power smoothing engages.
Companies: NVIDIA