NVIDIA CEO Jensen Huang Leaves Everyone SPEECHLESS (CES Supercut)
By Ticker Symbol: YOU
Key Concepts
- Vera Rubin: NVIDIA’s latest generation GPU and CPU architecture, designed for accelerated AI computing.
- GB200/GB300: Previous and current generation GPU platforms leading up to Vera Rubin.
- Extreme Code Design: A holistic approach to chip design, innovating across all components of the system simultaneously to overcome limitations of Moore’s Law.
- Spatial Multi-Threading: A CPU core design technique maximizing thread performance by treating multiple threads as independent cores.
- MVFP4 Tensor Core: A dynamically adaptive floating-point processing unit within the GPU, optimizing precision and throughput for transformer models.
- Spectrum-X: NVIDIA’s AI-optimized Ethernet networking technology.
- BlueField 4: A Data Processing Unit (DPU) for offloading networking and security tasks, and providing expanded context memory.
- MGX Chassis: NVIDIA’s modular, standardized compute chassis for building AI systems.
- HBM (High Bandwidth Memory): Fast memory used in GPUs for storing model parameters and intermediate results.
- KV Cache: The memory storing past interactions in a language model, crucial for context.
The AI Compute Race and the Vera Rubin Architecture
The presentation centers on NVIDIA’s response to the rapidly escalating demands of Artificial Intelligence, specifically the need for exponentially increasing compute power. The core argument is that while Moore’s Law is slowing, the intensity of the AI race necessitates continuous, aggressive innovation across the entire computing stack – a strategy NVIDIA terms “extreme code design.” The speaker emphasizes that the cost of AI tokens is decreasing by a factor of 10x annually, driven by competition and advancements in computing. This necessitates a constant push for faster computation to reach the “next frontier” of AI capabilities.
Vera Rubin: A Revolutionary System
NVIDIA’s latest offering, the Vera Rubin platform, is presented as a complete overhaul of their compute architecture. Recognizing the limitations of simply increasing transistor counts, NVIDIA redesigned every chip in the system. This includes:
- Vera CPU: A power-constrained CPU delivering 2x the performance per watt of the leading competitors, with an “insane” data rate. It features 88 physical cores utilizing “spatial multi-threading,” effectively functioning as 176 cores, each capable of full performance.
- Reuben GPU: A GPU that significantly increases single-threaded performance and memory capacity. While only 1.6x the transistor count of Blackwell (the previous generation), it achieves a 5x increase in floating-point performance.
- MVFP4 Tensor Core: This is a key innovation. Unlike traditional FP4 or FP8 implementations, the MVFP4 is a full processor dynamically adjusting precision to maximize throughput while maintaining accuracy. NVIDIA has published papers on this technology and anticipates it becoming an industry standard.
- MGX Chassis Redesign: The chassis itself has been radically redesigned, reducing assembly time from 2 hours to 5 minutes and achieving 80-100% liquid cooling.
- Networking – Spectrum-X & BlueField 4: NVIDIA’s Spectrum-X AI Ethernet is highlighted as the world’s best networking solution for high-performance computing, delivering 25% higher throughput. BlueField 4, a new Data Processing Unit (DPU), offloads networking and security tasks and provides a massive expansion of context memory (KV cache) for AI models.
- MVLink & Copper Interconnects: The system utilizes MVLink, with 2 miles of shielded copper cables (5,000 cables total) to achieve incredibly high bandwidth between components.
Technical Specifications and Performance Gains
The presentation details several key performance metrics:
- Transistor Count: Vera Rubin has 1.7x more transistors than Blackwell.
- Peak Inference Performance: 5x higher than Blackwell.
- Peak Training Performance: 3.5x higher than Blackwell.
- Energy Efficiency: Despite doubling power consumption, Vera Rubin maintains similar airflow and water cooling requirements (45°C water temperature). This translates to a 6% reduction in global data center power consumption.
- Power Smoothing: Vera Rubin incorporates power smoothing technology to address the high current spikes associated with AI workloads, eliminating the need for over-provisioning.
- Training Throughput: Vera Rubin enables training of a 10 trillion parameter model on 100 trillion tokens in one month, compared to four months with Blackwell.
- Factory Throughput: 10x higher than Blackwell.
- Cost of Tokens: 1/10th the cost of tokens compared to previous generations.
The Importance of Networking and Memory
The speaker emphasizes the critical role of networking and memory in AI performance. NVIDIA’s acquisition of Mellanox and the development of Spectrum-X demonstrate their commitment to optimizing network performance. The introduction of BlueField 4 and its associated 150TB of context memory per GPU (adding 16TB per GPU on top of existing memory) addresses the growing need for larger context windows in AI models. The MVLink 72 switch, with its 400 Gbits/second certis, enables every GPU to communicate with every other GPU simultaneously, achieving a bandwidth equivalent to twice the global internet data rate.
Standardization and Ecosystem
NVIDIA is promoting standardization through the MGX system, with approximately 80,000 components. This allows major computer manufacturers like Foxconn, Quanta, HP, and Dell to easily build and deploy NVIDIA-based AI systems.
Notable Quotes
- “The faster you compute, the sooner you can get to the next level of the next frontier.”
- “If we don't do code design, if we don't do extreme code design at the level of basically every single chip across the entire system, how is it possible we deliver performance levels that is, you know, at best one point 1 1.6 times each year?”
- “This is the world’s first manufacturing chip using TSMC’s new process that we co-innovated called coupe.”
Synthesis and Conclusion
The presentation paints a picture of NVIDIA as a full-stack AI company relentlessly pursuing innovation to overcome the limitations of traditional scaling methods. Vera Rubin represents a significant leap forward in AI compute capabilities, driven by “extreme code design” and a holistic approach to system architecture. The focus on energy efficiency, networking, and expanded memory capacity positions NVIDIA to continue leading the AI revolution and enabling the development of increasingly powerful and cost-effective AI applications. The core takeaway is that continued progress in AI requires not just faster chips, but a complete reimagining of the entire computing infrastructure.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

The Unheard-Of A+ Stock: Why This Tech Pullback is a Golden Opportunity
Seeking Alpha

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial