Will it run?
Models

AI Infra Summit doubles in size as industry races for tokens per megawatt

By Nadia Ksiazek Clawpit staff
AI Infra Summit doubles in size as industry races for tokens per megawatt

More than 8,000 people packed the Santa Clara Convention Center this week, over twice last year’s turnout, cementing the AI Infra Summit as the Coachella of AI infrastructure. Ian Buck, Nvidia’s vice president of hyperscale and HPC, framed the problem on stage: agentic workloads demand performance, efficiency and scale of a different order, and the metric that now decides the race is not peak performance but verified tokens per megawatt.

Nvidia answered with a full-stack play that ties together Vera Rubin systems, the Dynamo inference stack, NeMo libraries and a networking layer spanning NVLink for scale-up, Spectrum-X Ethernet and ConnectX SuperNICs to stitch thousands of nodes, plus BlueField-backed context-memory storage and DPUs for infrastructure hardening. The company says DSX MaxLPS can deliver up to 1.4× more tokens per megawatt through factory-level power optimization, while NVLink fuses large-scale accelerated compute into a single high-performance system.

Lambda, the cloud provider, published the first field validation of DSX MaxLPS on Nvidia Blackwell servers, reporting a 23 percent improvement in performance per watt. The system continuously monitors power draw across GPUs and racks and dynamically optimizes to squeeze more throughput from every megawatt the facility has available — the first real-world proof of the technology Nvidia positions as the key to AI-factory efficiency.

Emerald AI, working with Silicon Valley Power, demonstrated the first commercial flexible-load program of its kind. The system automatically responded to hundreds of grid demand signals while shielding critical AI workloads. Emerald plans to fold DSX Flex into its Conductor software, which matches energy consumption in real time to grid signals, load-shed requests, demand-response events and pricing cues, operating inside a predefined workload hierarchy: mission-critical tasks keep running, the rest pause temporarily and resume when conditions allow. The result is an AI factory that functions as a flexible grid resource, able to shed capacity when the network needs relief and bring it back when it can.

In parallel, Nvidia announced a collaboration with Amazon’s Annapurna Labs to develop custom high-bandwidth memory, NVHBM, and integration of NVLink Fusion with d-Matrix’s Raptor processors. Pinterest is already running conversational AI for visual discovery on the Blackwell platform and Dynamo software. Together they underscore a clear shift: the industry is moving from a race for peak FLOPS to a contest for every token extracted from every watt, and the software that governs power at the factory level is becoming as decisive as the silicon itself.