Will it run?
Agents

Nvidia says Vera Rubin NVL72 delivers 30× watt output over GB300 NVL72 on agent workloads

By Desmond Okafor Clawpit staff
Nvidia says Vera Rubin NVL72 delivers 30× watt output over GB300 NVL72 on agent workloads

OpenRouter data indicate that AI agent workloads consume roughly 15 times more tokens than a simple chat request. The increase stems from an agent that, for an investment decision, retrieves financial databases, scans news and reports, runs a sub-agent for peer comparisons and valuation, and then synthesizes a recommendation; tokens generated at each step become input for the next, making long-context handling critical. The same pattern appears in software development, customer-service automation, and deep research.

Nvidia measured performance with the SemiAnalysis AgentX workload, which uses real-world sessions of agent coding that grow context length, invoke tool calls, and activate sandboxed sub-agents. Although the results await SemiAnalysis review, they show that the GB300 NVL72 already provides about 15 times the megawatt-per-output of the Hopper architecture when running the DeepSeek V4 Pro model. Vera Rubin extends that advantage to roughly 30 times the output of GB300 in the same model. The figures do not yet include Vera CPU performance for tool calls.

The tests covered the Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro models. Nvidia’s DSX MaxLPS technology manages power at the GPU, rack and workload levels, allowing up to 40 % more GPUs within the same megawatt budget and pushing watt output higher at AI-factory scale. Because megawatt output drives revenue and cost per million tokens drives margin, the improvement translates to up to 35 times lower cost per million tokens compared with GB300 NVL72.

For power-constrained AI factories, the numbers suggest that agents can be run continuously at scale across diverse customer workloads. Nvidia notes that successive software optimizations will keep improving performance in both Vera Rubin and GB300. The remaining questions concern how much of the current gap will close after the independent review is completed and when central-processor performance is incorporated into the analysis.