Nvidia says its next-generation Vera Rubin NVL72 platform can deliver up to 30 times more AI throughput per megawatt compared with the prior generation, along with as much as 45 times lower token costs, according to figures the company published this week alongside benchmark data from SemiAnalysis's AgentX suite. The platform is now in full production.

The announcement leans heavily on power efficiency rather than raw chip supply, reflecting a shift in what is now constraining large AI data centers. Nvidia and outside analysts increasingly point to electricity availability, not chip manufacturing capacity, as the binding limit on how fast AI infrastructure can scale, particularly as agentic AI workloads that chain together many model calls put sustained strain on existing facilities.

Vera Rubin NVL72 is built around a seven-chip ecosystem and supports techniques such as disaggregated serving, expert parallelism and distributed key-value caching, all aimed at squeezing more useful AI work out of a fixed power budget rather than simply adding more hardware.

AdvertisementIn-Article

Source: NVIDIA