Author: Evolving AI - Bewertung: 0x - Views:1
NVIDIA just announced the moment every AI lab has been waiting for: Rubin, a next-generation GPU platform that doesn’t just beat Blackwell, it nukes the cost of running frontier models. In this video, we break down the full Vera-Rubin system: six chips acting as one AI supercomputer, with the Vera CPU (227 billion transistors and 88 Olympus cores), Rubin GPUs delivering up to 50 PFLOPS per pod and 100 PFLOPS in the NVL72 rack, 288 GB of HBM4 at 22 TB/s per GPU, and a networking stack built from ConnectX-9, BlueField-4 and Spectrum-X photonic Ethernet. NVIDIA is claiming up to 10× lower inference token costs vs Blackwell, 4× fewer GPUs for MoE training, 5× better energy efficiency, and a full re-think of cost per token through new NVFP4/6/8 precisions and aggressive HBM4 packaging. This isn’t a small generational bump – it’s an attempt to rewrite the economics of trillion-parameter AI. We also unpack how Rubin turns data movement into the real battleground, from NVLink 6 racks that let 72 GPUs behave like one training organism, to Spectrum-X photonic routing that treats Ethernet like a high-bandwidth extension of the GPU instead of a bottleneck. And then we zoom out to Alpamayo, NVIDIA’s open-source physical AI stack for cars and robots, and how its 10B-parameter Vision-Language-Action model takes a totally different path from Tesla FSD by reasoning before acting and exposing its chain of thought for regulators and OEMs. Together, Rubin and Alpamayo are NVIDIA’s answer to the next decade of AI: cheaper inference at insane scale, and physical AI that can explain itself. And if you want the real story behind the world’s fastest-moving AI breakthroughs, make sure to like and subscribe to Evolving AI for daily coverage.
SOCIAL SHARE CARD GENERATOR