Will it run?
Models

Nvidia expands Vera Rubin with Grok 3 LPX in full production, hits 3,400 tokens a second in long-context test

By Rae Whitlock Clawpit staff
Nvidia expands Vera Rubin with Grok 3 LPX in full production, hits 3,400 tokens a second in long-context test

In an Artificial Analysis benchmark running Gemma 4 31B — an agentic model with open weights — the system delivered 3,400 output tokens per second across a context window of 100 thousand tokens, four times the rate of the nearest competing platform. The figure applies specifically to long-context inference scenarios that multi-agent systems depend on, not to generic throughput.

Nebius, a major AI cloud provider, is the first to adopt Grok 3 LPX and will integrate it into its Nebius Token Factory alongside Vera Rubin NVL72 racks. CoreWeave has already deployed Spectrum-X Multiplane in production, a network architecture that links Vera Rubin racks through parallel switches to create a flat, lossless, high-bandwidth fabric. SpaceXAI said it will build its future AI architecture on Vera processors, from ground data centers to satellites in orbit.

The announcement came this week at the Hot Chips conference in Palo Alto. Nvidia is framing the stack as "extreme codesign": compute, networking, and inference acceleration engineered as a single system rather than as discrete components. Spectrum-X Ethernet handles the massive inter-rack data flow, Grok 3 LPX targets token generation at ultra-low latency, and the Vera Rubin NVL72 serves as what the company calls its most flexible factory platform.

The shift from training to inference and to agentic systems creates fundamentally different demands: token volumes are growing, context windows are expanding dramatically, and agents collaborate to solve complex problems. The required infrastructure is measured not only in peak performance but in throughput, responsiveness, and economics at unprecedented scale. Nvidia is positioning the platform as an integrated engine that turns ever-larger token volumes into revenue — phrasing that underscores the business logic behind the architecture.