Nvidia Rubin GPU: AI Agentic Throughput & Record Performance
Summary
Nvidia has unveiled its new Rubin GPU architecture and Vera CPU with the Olympus core. This comes alongside significant updates to its AI platforms and tools. What's new is that Nvidia's GB300 NVL72 system set a world record for mixture-of-experts pre-training performance. It achieved 1,648 teraflops per GPU when pre-training the DeepSeek-V3 671B model using 256 GPUs. This is roughly three times the per-GPU throughput of the earlier GB200 NVL72. The Rubin GPU is designed to deliver up to 10 times more agentic throughput per unit of energy than Blackwell. It contains 336 billion transistors and 896 Tensor Cores. It also integrates up to 288 GB of HBM4 memory with up to 22 terabytes per second of peak bandwidth. Software optimizations also show major gains. TorchTitan achieved roughly six times higher performance through Nvidia optimizations, and JAX performance improved by nearly 10 times over six months. The bottom line is these advancements aim to boost the efficiency and power of AI workloads, impacting future technological developments.
This is an AI-generated audio summary. Always check the original source for complete reporting.