Nvidia has a new headline number, and it is a big one. The company says its Vera Rubin NVL72 systems deliver up to 30 times more AI-factory throughput per megawatt than the GB300 generation on demanding agentic workloads, along with a claimed 45 times lower token cost. Those figures, published on the Nvidia blog, describe the second generation of the company's rack-scale Oberon architecture and its bet that the future of inference is agents running long, tool-heavy tasks rather than single quick answers.
A 30x claim invites a careful reading, and the independent analysis supplies one. In a detailed comparison, SemiAnalysis found that on standard inference the advantage is nearer 2 times the throughput of the prior Blackwell racks up to around 100 tokens per second per user, widening to roughly 4 times near 200 tokens per second per user, where the gap peaks. Running DeepSeek R1, the firm measured about 5.4 times the performance per megawatt and 5 times the performance per dollar over GB200.
Two numbers, one honest gap
The two accounts are not in conflict so much as measuring different things. The eye-catching 30x figure comes from agentic benchmarks, where the workload stresses memory and coordination in ways that flatter the newer design. The more sober 2x to 5x range describes the everyday inference most operators actually run today. Both can be true. The useful takeaway is that Rubin is a genuine step up, and the size of the step depends heavily on what you ask it to do.
That distinction matters because the economics of a data center turn on total cost of ownership, not peak benchmarks. SemiAnalysis has cautioned that the case for Rubin narrows once power, cooling, and the price of the racks themselves are folded in. Efficiency per watt is only worth so much if the hardware costs proportionally more to buy and house.
Why the watt is the unit that counts
The steady move toward measuring AI in work-per-watt is telling. Compute is increasingly constrained by electricity rather than by chips on a shelf, which is why grid capacity has become a live political issue and why challengers are attacking Nvidia from other angles, including the interconnect. A startup like Cornelis is arguing that the network should shoulder part of the computing, and the custom-silicon deals reshaping the supplier landscape, such as Broadcom's surging AI chip business, point the same way.
For now Nvidia still sets the pace. The Vera Rubin numbers are impressive on their own terms, and the caveats do not erase them. They just put the marketing where it belongs, next to the benchmark that produced it.
Commentarii · 0