AI · Sep 16, 2026
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 DebutAI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
Sep 15, 2026, 9:55 AM · NVIDIA Blog

At AI Infra Summit, NVIDIA’s Ian Buck highlighted Vera Rubin, DSX, Emerald AI flex load with Silicon Valley Power, Lambda’s 23% perf-per-watt gain, and partner deals from Annapurna to d-Matrix.
Why it matters
NVIDIA reports Ian Buck speaking to 8,000+ AI Infra Summit attendees (up from 3,500) on factory efficiency: Amazon Annapurna collaborating on NVHBM custom HBM; d-Matrix integrating NVLink Fusion with Vera CPUs and Raptor XPUs for low-latency inference; Emerald AI flexible-load demo with Silicon Valley Power; Lambda claiming 23% better performance per watt via DSX MaxLPS; Pinterest on Blackwell plus Dynamo for conversational visual discovery.
Agentic workloads are the demand shock. Tokens per watt is the answer NVIDIA wants on every slide.
Partner collage events matter when the metrics are checkable.
From the desk
We’re scanning for numbers that survive outside the keynote hall.
Lambda’s 23% perf-per-watt with DSX MaxLPS is the kind of claim operators can try to reproduce. Emerald AI’s utility flex ties efficiency to civic license. Annapurna NVHBM and d-Matrix NVLink Fusion show NVIDIA threading custom silicon into its fabric rather than fighting every accelerator alone.
Useful AI infra is boringly efficient. We’re for software that recovers stranded power and partners that cut latency for real products like Pinterest’s discovery. The harm of summit theater: CapEx decisions on unreplicated percentages, and “efficiency” talk that never reduces absolute community load.
I’m watching Lambda’s 23% methodology and whether NVHBM/NVLink Fusion partner timelines slip.
Context
AI Infra Summit in Santa Clara featured NVIDIA VP Ian Buck. DSX is NVIDIA’s data-center software stack for power and production optimization; Dynamo is inference software.
Who feels it
- Neoclouds like Lambda
- Public perf-per-watt gains become competitive marketing—customers should ask for workload-matched proofs.
- Custom-silicon builders
- NVLink Fusion path may beat going fully alone on scale-up fabric.
- Consumer platforms
- Pinterest-style agentic features depend on inference stacks, not only model weights.
What to watch
- Technical detail behind Lambda’s 23% DSX MaxLPS result.
- NVHBM sampling/production dates with Annapurna.
- Customer deployments of Vera CPU + d-Matrix Raptor via NVLink Fusion.
Companies: NVIDIA