2 min read
Add as a preferred source on Google

NVIDIA Introduces Multi-GPU Solver for 100M-Variable Models

The cuOpt solver uses NVLink-connected GPUs to handle models with up to 2.1 billion nonzero matrix entries while cutting per-GPU memory use by up to 6x.

Graphics cards connected inside an open computing test chassis / TokenPost.ai
Graphics cards connected inside an open computing test chassis / TokenPost.ai

NVIDIA introduced a multi-GPU solver for linear-programming models with more than 100 million variables, bringing distributed GPU computing to large supply-chain and energy-planning workloads in its cuOpt platform.

The Multi-GPU Primal-Dual hybrid gradient for Linear Programming (mPDLP) solver divides sparse matrix calculations across NVLink-connected GPUs. It is designed for problems with as many as 2.1 billion nonzero matrix entries and can reduce peak memory use per GPU by up to 6x compared with single-GPU PDLP.

Performance gains increased with workload size. Tests across more than 100 instances on NVIDIA DGX B200 systems showed noticeable speedups above 10 million nonzero entries. On the tsp-gaia-10m benchmark, mPDLP delivered up to an 11.4x improvement in PDLP-step performance and a 4.2x improvement in total processing time.

The smaller end-to-end gain reflected preprocessing and postprocessing steps that were not distributed across GPUs. On most large instances with more than 10 million nonzero entries, mPDLP produced a 1.2x-to-2.5x speedup over the earlier D-PDLP method. It was slower than D-PDLP on three ultra-large benchmarks, showing that results vary with model size and sparsity.

The solver also improved performance on industry workloads. A consumer-packaged-goods supply-chain model from Kinaxis with more than 135 million variables ran 3.3x faster using eight NVLink-connected NVIDIA H100 GPUs. A stochastic energy-expansion model from PSR achieved more than a fivefold speedup with eight NVIDIA B200 GPUs.

Linear programming helps organizations choose among constrained options, such as how to plan supply networks or expand energy systems. PDLP is a first-order optimization method built for GPU execution, while mPDLP uses min-cut partitioning to divide the workload and limit communication between processors.

The distributed-solver tests targeted a tolerance of 10⁻⁶ and limited each run to one hour. Further development is focused on load-aware partitioning, overlapping communication with computation and feasibility-polishing features.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…