SDSignal Desk

d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

Sep 10, 2026, 6:00 AM · NVIDIA Blog

Image: NVIDIA Blog

Inference chipmaker d-Matrix is plugging next-gen Raptor XPUs into NVIDIA’s NVLink Fusion stack — betting that rack-scale deployment, not just silicon, is the bottleneck.

Why it matters

d-Matrix said it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure — NVLink scale-up, Spectrum-X scale-out, MGX racks, and the broader platform. The company pitches faster, lower-risk paths to ultralow-latency inference under tight capital, time, and energy constraints.

Building a custom XPU is hard; shipping it at AI-factory scale is harder. Fusion is NVIDIA’s answer for third-party XPUs and CPUs: keep the partner’s differentiated silicon, reuse NVIDIA’s interconnect, rack designs, networking, cooling, and supply chain.

d-Matrix plans single-domain NVLink scale-up for its XPUs and coexistence with GPU systems such as Vera Rubin NVL72 for disaggregated inference, plus Vera CPUs, ConnectX-9, BlueField-4, and Spectrum-X in the same factories.

From the desk

We’re reading this as infrastructure politics as much as a chip announcement. NVIDIA is opening the factory door just enough that specialized inference silicon can live inside its rack and networking standard — which accelerates partners like d-Matrix and also deepens NVIDIA’s role as the default deployment substrate.

That’s useful when it shortens time-to-rack for architectures tuned to inference latency and efficiency. Customers get fungible halls that can mix GPUs, CPUs, and XPUs without a bespoke mechanical and thermal redesign per vendor. d-Matrix’s CEO is explicit: demand is soaring, resources aren’t, so riding MGX and Fusion beats reinventing liquid-cooled scale-up. For an inference-focused XPU house, borrowing a proven supply chain and cooling design may matter more than another percent on a die shootout.

The downside if this becomes the only practical path: semi-custom “openness” that still orbits one company’s networking and software assumptions. Partners may move faster and still find switching costs waiting at the rack boundary. And press-brief numbers on NVLink Fusion performance — lower XPU-to-XPU latency than off-the-shelf Ethernet, higher packet rates, terabytes-per-second class all-to-all bandwidth — are platform claims; d-Matrix’s own Raptor results at scale aren’t in this post.

I’m watching whether Fusion partners ship customer-visible racks on a calendar, not just logos on a slide — and whether disaggregated inference beside Rubin-class systems shows up in production designs. Until those land, this is a credible deployment alliance, not yet a measured inference win.

Context

NVIDIA lists a wide Fusion ecosystem spanning cloud, CPU, and chiplet partners — including names such as AWS, Arm, Intel, Fujitsu, and others cited in the same brief. The platform pitch is a semi-custom AI factory: NVIDIA infrastructure plus partner XPU, aimed at matching compute to workload inside one operational model that spans networking, storage, security, power, cooling, and software.

Who feels it

Inference chip startups
A faster deployment path — and strategic dependence on NVIDIA’s rack and networking stack.
Cloud and neo-cloud operators
More options to mix specialized XPUs into existing MGX-style designs if validation holds.
Enterprise AI buyers
Potential for workload-matched inference silicon without waiting on a fully custom factory build.

What to watch

  1. Concrete Raptor-on-Fusion timeline and customer deployments.
  2. How d-Matrix racks interoperate in practice with Vera Rubin NVL72 for disaggregated inference.
  3. Whether more inference specialists announce Fusion adoption on similar terms.

Read the original

Continue at the source.

NVIDIA Blog

Companies: NVIDIA