d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
Sep 10, 2026, 6:00 AM · NVIDIA Blog

Inference chipmaker d-Matrix is plugging next-gen Raptor XPUs into NVIDIA’s NVLink Fusion stack — betting that rack-scale deployment, not just silicon, is the bottleneck.
Why it matters
d-Matrix said it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure — NVLink scale-up, Spectrum-X scale-out, MGX racks, and the broader platform. The company pitches faster, lower-risk paths to ultralow-latency inference under tight capital, time, and energy constraints.
Building a custom XPU is hard; shipping it at AI-factory scale is harder. Fusion is NVIDIA’s answer for third-party XPUs and CPUs: keep the partner’s differentiated silicon, reuse NVIDIA’s interconnect, rack designs, networking, cooling, and supply chain.
d-Matrix plans single-domain NVLink scale-up for its XPUs and coexistence with GPU systems such as Vera Rubin NVL72 for disaggregated inference, plus Vera CPUs, ConnectX-9, BlueField-4, and Spectrum-X in the same factories.
From the desk
We’re reading this as infrastructure politics as much as a chip announcement. NVIDIA is opening the factory door just enough that specialized inference silicon can live inside its rack and networking standard — which accelerates partners like d-Matrix and also deepens NVIDIA’s role as the default deployment substrate.
That’s useful when it shortens time-to-rack for architectures tuned to inference latency and efficiency. Customers get fungible halls that can mix GPUs, CPUs, and XPUs without a bespoke mechanical and thermal redesign per vendor. d-Matrix’s CEO is explicit: demand is soaring, resources aren’t, so riding MGX and Fusion beats reinventing liquid-cooled scale-up. For an inference-focused XPU house, borrowing a proven supply chain and cooling design may matter more than another percent on a die shootout.
The downside if this becomes the only practical path: semi-custom “openness” that still orbits one company’s networking and software assumptions. Partners may move faster and still find switching costs waiting at the rack boundary. And press-brief numbers on NVLink Fusion performance — lower XPU-to-XPU latency than off-the-shelf Ethernet, higher packet rates, terabytes-per-second class all-to-all bandwidth — are platform claims; d-Matrix’s own Raptor results at scale aren’t in this post.
I’m watching whether Fusion partners ship customer-visible racks on a calendar, not just logos on a slide — and whether disaggregated inference beside Rubin-class systems shows up in production designs. Until those land, this is a credible deployment alliance, not yet a measured inference win.
Context
NVIDIA lists a wide Fusion ecosystem spanning cloud, CPU, and chiplet partners — including names such as AWS, Arm, Intel, Fujitsu, and others cited in the same brief. The platform pitch is a semi-custom AI factory: NVIDIA infrastructure plus partner XPU, aimed at matching compute to workload inside one operational model that spans networking, storage, security, power, cooling, and software.
Who feels it
- Inference chip startups
- A faster deployment path — and strategic dependence on NVIDIA’s rack and networking stack.
- Cloud and neo-cloud operators
- More options to mix specialized XPUs into existing MGX-style designs if validation holds.
- Enterprise AI buyers
- Potential for workload-matched inference silicon without waiting on a fully custom factory build.
What to watch
- Concrete Raptor-on-Fusion timeline and customer deployments.
- How d-Matrix racks interoperate in practice with Vera Rubin NVL72 for disaggregated inference.
- Whether more inference specialists announce Fusion adoption on similar terms.
Companies: NVIDIA