How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency
Sep 27, 2026, 6:00 PM · NVIDIA Developer

A joint NVIDIA–Nscale GB300 run shows policy-shared power packing ~37% more GPUs and ~49% more tokens under the same 264.4 kW budget—with a P99 TTFT tax.
Why it matters
AI factories still provision for the rare moment every GPU spikes at once. That buffer keeps the site safe—and leaves watts idle while racks sit dark. NVIDIA DSX MaxLPS uses Dynamic Power Software to share power across a managed group under operator policy, claiming up to 40% more GPUs inside the same approved budget.
A joint evaluation with Nscale at the Verne campus in Keflavík ran Kimi K2.5 on GB300 NVL72 systems: static baseline 140 GPUs versus MaxLPS at 192 GPUs, same 264.4 kW provisioned power. Aggregate throughput rose 49.2%; throughput per provisioned watt moved from 4.10 to 6.12 tokens/s/W. Per-instance throughput for high-throughput and low-latency jobs stayed effectively flat.
Power is the binding constraint for useful AI at factory scale. Reclaiming stranded headroom without lying about latency is the operator problem this piece actually measures.
From the desk
We’re taking the measured trade-off seriously, not the brochure ceiling. Going from 140 to 192 GPUs (+37.1%) under one budget, with aggregate tokens/s up ~49% and mean GPU power up ~36%, is the story of complementary workloads funding an extra high-throughput instance. Per-instance rates barely moved—so the gain is fleet packing, not magic per chip.
Median and P75 latency stayed within 5% of baseline. P99 time to first token rose 17% from a 15.7-second baseline. That’s the bill: capacity and average service can look fine while the slowest percentile stretches. Anyone selling “more GPUs, same power” without publishing tail numbers is skipping the part that hurts interactive users.
Useful AI factories should chase tokens under real electrical limits. MaxLPS doesn’t add utility feeders; it reallocates unused reservation. Workload mix decides how much headroom exists—homogeneous peaks leave less to share. Telemetry integrity is a hard dependency: bad maps or delayed power readings make the control loop a liability.
NVIDIA’s five-stage validation path—boundary, baseline, conservative policy, incremental capacity, production limits—is the part operators should steal even if they never buy the brand name. Plan cooling and fabric for the validated target, not day-one rack count. And keep Vera Rubin NVL72 projections separate from this GB300 measurement, as the post itself warns. I’m watching for third-party sites to publish the same protocol with their own mix—and whether that 17% P99 move is acceptable for their SLOs.
Context
NVIDIA developer blog by Sarah McKenney and Harry Petty, September 27, 2026. Evaluation: Nscale Verne campus (Keflavík), renewable-powered; Blackwell Ultra / GB300 NVL72; Kimi K2.5 FP4; Dynamo and TensorRT LLM; 8K in / 1K out; jobs confined to single racks across four racks.
Who feels it
- AI factory operators
- Same provisioned watts can host more GPUs if policy, telemetry, and acceptance criteria include P99—not only aggregate tokens/s.
- Neoclouds and colos
- Power-sharing becomes a commercial lever; electrical, cooling, and fabric still must be sized for the validated lifecycle target.
- Inference product owners
- Fleet gains that hold median latency can still move TTFT tails—set SLOs before celebrating utilization.
What to watch
- Independent MaxLPS-style results outside the NVIDIA–Nscale joint study
- Whether production SLOs explicitly cap P99 TTFT when packing denser under fixed power
- How Vera Rubin NVL72 dynamic-power claims compare once measured—not projected—against this GB300 baseline
Companies: NVIDIA