SDSignal Desk

Scaling Decision Optimization to 100 Million Variables and Beyond with mPDLP in NVIDIA cuOpt

Oct 7, 2026, 8:45 AM · NVIDIA Developer

Image: NVIDIA Developer

NVIDIA's multi-GPU solver makes huge planning problems for supply chains and power grids faster to solve, though only past a certain size, and its own benchmarks show where it loses.

Why it matters

NVIDIA has added a multi-GPU linear programming solver, mPDLP, to its cuOpt optimization library. It splits very large problems across NVLink-connected GPUs, which NVIDIA says cuts solve time on problems too big or slow for one GPU and lowers peak memory per GPU by up to 6x compared with its single-GPU solver.

This is not chatbot AI, but it is the math behind real decisions: how much to produce, where to ship, which power plants to build. Two partners reported results. Kinaxis saw a 3.3x speedup on a consumer-goods supply chain model with over 135 million variables using eight H100 GPUs, and PSR reported more than 5x on an energy expansion model with 185 million variables using eight B200 GPUs.

From the desk

We think this matters more than it sounds. When the largest planning problems take hours to converge, teams can end up testing fewer scenarios than they should. Faster solves on bigger models mean more uncertainty can be tested before a supply plan or a grid investment gets locked in. That is useful AI-adjacent infrastructure in the plainest sense.

What we appreciate is that NVIDIA shows the limits. Speedups only show up once a problem has more than about ten million nonzero entries; below that, the cost of keeping GPUs in sync outweighs the gain. Against an earlier multi-GPU method, D-PDLP, mPDLP was 1.2x to 2.5x faster on most large instances but slower on the three largest ones, which NVIDIA attributes, tentatively, to sparsity patterns that differ from the problems it tuned for. End-to-end, the headline 11.4x figure for the core solver steps on one benchmark shrinks to 4.2x once setup and cleanup are counted.

The downside is access. The benefits accrue to organizations that can run eight top-end GPUs on a single planning model. That favors large enterprises and vendors, and deepens dependence on one hardware stack for decisions that used to run on CPUs. I'm watching for independent benchmarks from optimization researchers, and for whether the load-balancing improvements NVIDIA lists close the gap on the hardest problems.

Context

The method is a GPU-friendly, first-order approach to linear programming called PDLP. NVIDIA's version partitions the problem so that tightly linked parts stay on the same GPU, reducing cross-GPU traffic. NVIDIA says it is working on load-aware partitioning, overlapping communication with computation, and faster feasible solutions. cuOpt is available on GitHub with a tutorial for the new solver.

Who feels it

Supply chain and energy planners
Larger, more detailed models become practical within planning windows, allowing more scenarios before decisions are made.
Optimization software vendors
Partners like Kinaxis and PSR are already integrating multi-GPU solving, raising the bar for competitors.
Smaller organizations
The gains require multi-GPU hardware, so the benefit is concentrated among those who can afford it.

What to watch

  1. Independent benchmarks against leading commercial and open-source LP solvers
  2. Results from the planned load-aware partitioning and communication overlap
  3. More optimization platforms adopting mPDLP
  4. Performance on the ultra-large instances where it currently trails D-PDLP

Read the original

Continue at the source.

NVIDIA Developer

Companies: NVIDIA