E21: NVIDIA'S HUGE AI Chip Breakthroughs Change Everything
By Ticker Symbol: YOU
Key Concepts
- AI Training: The initial phase of teaching foundational knowledge to AI models.
- Post-Training: Injecting specialized, industry-specific knowledge into a foundational AI model.
- Inference: The phase where a trained and fine-tuned AI model is put to use to generate predictions or responses for end-users.
- Prefill: The compute-heavy initial stage of inference where the AI model processes and understands the context of a user's query or prompt.
- Decode: The memory-latency-bound subsequent stage of inference where the AI model auto-regressively predicts and generates tokens sequentially.
- Foundational Models: Large, general-purpose AI models trained on vast datasets, forming the basis for more specialized applications (e.g., ChatGPT, Claude).
- Large Language Models (LLMs): A type of foundational model specialized in understanding and generating human language.
- High Bandwidth Memory (HBM): Specialized memory crucial for the decode phase of inference due to its high speed and low latency.
- Rubin CPX GPU: Nvidia's purpose-built GPU optimized for "million context workloads" (prefill-heavy inference).
- Rubin Platform: Nvidia's general-purpose platform, highly efficient for the decode phase of inference.
- Extreme Co-Design: Nvidia's holistic approach to system design, integrating GPUs, CPUs, DPUs, and networking components, along with software, to achieve massive performance leaps beyond traditional Moore's Law gains.
- CPU (Central Processing Unit): The general-purpose processor, co-designed with GPUs for tightly coupled memory and processing.
- DPU (Data Processing Unit - BlueField): A specialized processor for data movement, storage access, and node interaction within the AI infrastructure.
- NVLink Switch Technology: Nvidia's custom chip for "scale-up" architecture, allowing multiple GPUs, CPUs, and processors to function as a single, powerful unit.
- Spectrum X Ethernet Switching Technology & Infiniband Network: Nvidia's networking solutions for "scale-out" architecture, connecting hundreds of thousands to millions of GPUs.
- MVFP4: A specific precision format introduced with Blackwell, requiring software integration for intelligent utilization.
- Blackwell, Hopper, Blackwell Ultra, Rubin, Rubin CPX: Generations and specialized variants of Nvidia's GPU architectures and platforms.
- AI Factory: A business leveraging AI at scale to solve problems, transforming energy and components into intelligence.
- Cost per Token: A key metric for AI efficiency, representing the cost to generate a single unit of AI output.
- DSX (Digital Twin Simulation for AI Factories): A GTC-announced initiative providing digital twin capabilities and reference blueprints for building efficient gigawatt-scale AI factories.
- Moore's Law: The observation that the number of transistors in an integrated circuit doubles approximately every two years, which has tapered off as a primary driver of performance.
Nvidia's Comprehensive Approach to AI: Beyond Training to Inference
The discussion with Dion Harris, Nvidia's Senior Director of High Performance Computing, Cloud, and AI Infrastructure Go-to-Market, reveals that Nvidia's leadership in AI extends significantly beyond its widely recognized role in AI training. The company's strategy encompasses the entire lifecycle of AI, with a particular emphasis on the complex and rapidly evolving domain of inference.
Phases of AI Development and Nvidia's Infrastructure Support
Nvidia's GPUs and infrastructure support all critical phases of AI:
- AI Training: This initial phase involves teaching foundational models (e.g., ChatGPT, Claude) vast amounts of knowledge, often from the entire content of the internet. Training focuses on enabling models to learn semantic language and understand meanings across different modalities.
- Post-Training: After foundational training, models undergo post-training to inject specialized knowledge. This fine-tuning tailors a general model to specific industries or use cases. For example, a model might be post-trained to understand the nuanced terminology of healthcare, where a term like "cell" has a different meaning than in the legal profession.
- Inference: This is the phase where AI is put to practical use. Once a model is trained and fine-tuned, inference involves deploying it to interact with end-users, customers, or partners, thereby extracting business value. While early AI investment focused on building training clusters, the current emphasis is on deploying AI for inference at scale, marking a "critical tipping point" in AI's utility.
Deeper Dive into Inference: Prefill and Decode
Inference itself is not a monolithic process but can be decoupled into distinct workloads with varying computing requirements:
- Prefill: This is the compute-heavy phase where the AI model processes the context of a user's query. This includes understanding the prompt, uploaded documents (as source information), or previous interactions. The goal is to build "contextual awareness" by processing all input tokens.
- Example: A deep research project requiring the AI to sift through numerous uploaded PDFs would be highly prefill-heavy.
- Decode: Following prefill, the decode phase involves the auto-regressive prediction of each subsequent token to generate a response. This phase is memory latency bound, requiring high-bandwidth memory (HBM) to generate tokens rapidly.
- Example: Asking an AI to produce a long, in-depth codebase would be decode-heavy, as it requires extensive reasoning chains to generate high-quality, sequential code.
The balance between prefill and decode varies significantly based on the model, user profile, and specific request, making inference particularly challenging.
Nvidia's Specialized Hardware for Inference
To address these distinct requirements, Nvidia develops specialized hardware:
- Rubin CPX GPU: This GPU is purpose-built for "million context workloads," which are highly prefill-intensive. It provides immense compute power for tasks like advanced code generation (understanding entire application codebases) and video generation (maintaining contextual consistency across long videos). The design goal is to provide high compute while lowering memory costs, as prefill is less memory-intensive.
- Rubin Platform: Nvidia's standard Rubin platform is already highly efficient for the decode phase, meaning the CPX was developed to specifically optimize the prefill step that needed a dedicated solution.
Nvidia's Extreme Co-Design Strategy
Nvidia's ability to deliver generational leaps in AI performance stems from its "extreme co-design" approach, recognizing that Moore's Law has tapered off. This strategy involves co-designing an entire system, not just individual chips:
- Beyond the GPU: Performance is driven by the synergistic interaction of:
- CPU: Tightly coupled with the GPU for memory and processing power.
- DPU (BlueField): Manages data movement within GPUs, to/from storage, and node access, integrating with software.
- NVLink Switch Technology: A custom chip enabling "scale-up" architecture, allowing multiple GPUs, CPUs, and processors to function as a single, unified computing entity.
- Spectrum X Ethernet Switching Technology & Infiniband Network: Solutions for "scale-out" architecture, connecting hundreds of thousands to millions of GPUs across vast data centers.
- Software Integration: Co-design extends to the software layer and model developers. For instance, new precision formats like MVFP4 (introduced with Blackwell) require software to be intelligently taught how to leverage them. Nvidia actively works with the open-source community and developers to ensure their software takes full advantage of underlying hardware innovations.
- Annual Rhythm of Innovation: Due to the rapid evolution of AI models, Nvidia maintains an "annual rhythm" of platform development. This allows them to "leapfrog" themselves, continuously delivering more performance and value, as seen with the progression from Blackwell to Blackwell Ultra, and Rubin to Rubin CPX.
Business Value of AI Inference and Efficiency
The technical advancements in inference directly translate into significant business value:
- AI Factory Concept: Businesses leveraging AI at scale are essentially "AI factories," transforming energy and components into "intelligence." The key metric becomes "how much more intelligence can I produce per dollar or per watt."
- Efficiency as ROI Driver: Efficiency is the primary driver for return on AI investment. Improved inference performance per watt allows power-limited data centers to generate more intelligence within their existing power envelopes. These "X factors" directly translate into "actual dollars and cents" by enabling the generation of more tokens and extraction of greater value. For companies that monetize tokens, this directly correlates to increased revenue and profit.
- Forward-Looking Indicators: Cost per Token: A crucial metric for investors is the cost per token. As this cost decreases, AI can be embedded into more services and use cases, delivering greater value to end-users. A 10x reduction in cost per token can lead to a 20x increase in overall utilization, as previously unaffordable use cases become viable, expanding the "surface area of AI." This demonstrates a highly elastic demand for AI capabilities.
- Nvidia's True Advantage: While known for training, Nvidia believes its "true advantage lies in our ecosystem and our software maturity" for tackling the complexities of inference, making its platform even more valuable in inference than in training.
- Convergence of Training and Inference: Jensen Huang's prediction that training and inference will merge is already becoming a reality. Reasoning models are trained using extensive inference, creating an iterative feedback loop where inference outputs feed back into the training process, making the two "indistinguishable."
Keeping Up with Insane Demand and Future Outlook
To meet the projected "billionx" rise in inference demand over the next few years, Nvidia relies on its extreme co-design strategy:
- Exponential Performance Gains: The Blackwell platform demonstrated a 10x performance per watt improvement over the previous-generation Hopper in inference max results. This unprecedented leap was achieved not by simply adding more transistors, but by the comprehensive extreme co-design approach, including scale-up (NVL72 for parallelization), Dynamo software for disaggregated serving, and improved precision (MVFP4) while maintaining accuracy.
- Holistic Optimization: Nvidia optimizes across the entire stack—from individual chips to the tray, rack, data center, and even multiple data centers—to drive maximum efficiency.
- Modularity and Disaggregation: Despite building a fully integrated stack, Nvidia also designs its solutions to be modular and disaggregated. This allows users to integrate Nvidia components into their existing infrastructure, offering flexibility and catering to diverse business objectives (e.g., running Nvidia GPUs with other networking solutions or using NVLink Fusion for scale-up).
- DSX (Digital Twin Simulation for AI Factories): Announced at GTC, DSX provides digital twin capabilities and gigascale AI factory reference blueprints. This helps the entire ecosystem design, build, and operate highly efficient gigawatt-scale AI factories by leveraging Nvidia's best practices within digital environments.
- Developer-First Approach: Nvidia's excitement for the future stems from the collective effort across its partner and developer ecosystems. By creating conditions with new processors (like CPX), software (like Dynamo), and architectures, Nvidia aims to unlock entirely new sets of use cases, empowering developers to build innovations not yet conceived.
Conclusion
Nvidia's strategy for the future of AI is deeply rooted in its "extreme co-design" philosophy, extending its focus from AI training to the intricate and demanding world of inference. By optimizing every component of the AI infrastructure—from specialized GPUs like Rubin CPX to CPUs, DPUs, and advanced networking solutions—and integrating these with sophisticated software, Nvidia is achieving unprecedented performance gains (e.g., 10x perf/watt in a single generation). This relentless pursuit of efficiency drives down the "cost per token," making AI ubiquitous and unlocking immense business value by enabling a vast array of new applications and use cases. Through an annual rhythm of innovation, a developer-first approach, and tools like DSX, Nvidia is positioning itself at the epicenter of the AI transformation, ready to meet the exponential demand for intelligence.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
AI Engineer

'No where near normal' but 30-40 oil tankers passing through the Strait 'is better than 0': Mulberry
BNN Bloomberg

'Alphabet has such a dominant position they will be a leader in this space for many years': Clare
BNN Bloomberg

Forget Elon’s Data Centers In Space. This Startup Wants To Float Them At Sea
Forbes

Yahoo Finance Live: Daily Market Coverage - June 29, 2026 9AM-11AM (ET)
Yahoo Finance

GPT 5.6 banned, Fable banned… it’s actually over.
David Ondrej

GPT 5.6 Sol Just Blew Up The AI World
AI Revolution