CoreWeave launches Nvidia Vera Rubin NVL72; Cognition runs production workloads

At its Fully Connected conference in San Francisco, CoreWeave announced general availability of Nvidia Vera Rubin NVL72 systems equipped with Spectrum-X networking running at 102.4 terabits per second. Cognition, the lab behind the autonomous software engineer Devin, is the first customer running production workloads on the new platform. In an internal benchmark against a GB200 NVL72 baseline, the Vera Rubin system delivered up to 4.8 times total token throughput on SWE-2 workloads, which Cognition says translates to faster real-time code generation and expanded multi-step reasoning capability.
A decade-long partnership that keeps paying off
Nvidia's vice president of hyperscale and high-performance computing, Ian Buck, noted that the company's platforms continue to deliver value across generations: CoreWeave's V100 processors are still running customer workloads nearly a decade after the Volta architecture launched, even as the newest generation enters production. That flexibility — matching the right GPU to the right workload — is what lets CoreWeave keep selling capacity on older hardware.
Real benchmark, real code
Cognition scaled to thousands of GPUs on CoreWeave within nine months, using the cluster for training, reinforcement learning and production inference for Devin. When the first Vera Rubin NVL72 racks arrived this month, the company sampled tasks from FrontierCode and ran agents against them — a workload that mimics actual software engineering: long context, high concurrency and token volumes where cost per token determines what can be shipped. Silas Alberti, a member of Cognition's founding team, emphasized that the value for them lies not in any single spec but in a unified platform with Nvidia and CoreWeave engineers working alongside them on the hardest problems.
Forge, Vera CPU and sandboxes three times faster
Beyond the GPU, CoreWeave launched CoreWeave Forge, an integrated environment for training, evaluating and improving models and agents on Nvidia accelerated compute. Also coming to the cloud is the Nvidia Vera CPU, the first processor designed for agentic workloads. Early tests showed Vera improving agent sandbox spin-up time by more than three times. Agent workloads stress infrastructure from two directions: serving agents demands low-latency compute at scale, while improving them through post-training requires thousands of isolated environments running in parallel.
Deployment in days via Kubernetes and sandboxes
Capacity is available through CoreWeave Kubernetes Service, SUNK, Mission Control, Sandboxes and Inference. Thanks to codesign across the full stack — from infrastructure to served tokens — CoreWeave stood up a production Vera Rubin cluster for Cognition in a matter of days. Each rack contains 128 Vera CPUs and 11,264 cores, enough to run a large number of isolated agent environments simultaneously while maintaining consistent performance.