According to the NVIDIA official blog, compared to the GB300 NVL72, the Vera Rubin NVL72 system delivers 30 times higher throughput per megawatt and 35 times lower cost per million tokens when processing AI agent workloads. This performance improvement is crucial for power-constrained AI factories, meaning more agent work can be completed with the same energy consumption and significantly lower operating costs. NVIDIA measured these inference throughput figures using the SemiAnalysis AgentX workload, which includes real agentic coding sessions, preserving actual context growth, tool calls, and sub-agent generation.