Overview
- Nvidia and partners report that Vera Rubin — a liquid-cooled rack that pairs the new Vera CPU with Rubin GPUs — is in customer testing and has entered an initial volume production ramp this month.
- Vendor benchmarks and partner tests claim large efficiency gains, with some reports showing up to a 10x improvement in token throughput per megawatt and large reductions in inference token cost versus the prior Blackwell systems.
- Vera uses Nvidia’s Olympus 88-core monolithic CPU with a high-bandwidth on-die coherency fabric and LPDDR5X memory, and the NVL72 rack pairs 36 Vera CPUs with 72 Rubin GPUs over an all-to-all NVLink fabric.
- Independent analysts caution key inputs and infrastructure are constrained: HBM4 stacks, CoWoS advanced packaging, N3 wafer capacity and high-power datacenter cooling limit how quickly Nvidia can reach its cited peak output such as 1,000 racks per day.
- If Vera Rubin scales as claimed, cloud operators could cut inference costs and change CPU–GPU ratios in AI data centers, but hyperscalers and smaller customers will face practical hurdles from power, cooling and supply availability before those savings are widespread.