Conversation with NVIDIA CEO Jensen Huang: Satya Nadella at Microsoft Build 2025

MicrosoftAbout 4 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Moore's Law & its acceleration
  • AI Supercomputers on Azure
  • Full-stack innovation (hardware & software)
  • NVLink, Grace Blackwell, FP4 Tensor Core
  • CUDA, AI Algorithms, Model Technology
  • Fleet Optimization & Continuous Upgrades
  • Software Compatibility & Ecosystem Stability
  • Accelerated Compute & Workload Diversity
  • Tokens/Workloads per Dollar per Watt

Moore's Law and Accelerated Innovation

The discussion begins with the impact of Moore's Law and how NVIDIA is accelerating it. Satya Nadella highlights the goal of delivering more intelligence to the world, measured as "tokens per dollar per watt." Jensen Huang points out that two years prior, they launched the largest AI supercomputer on Azure, and now they are in full production with Grace Blackwell, scaling and building an even larger AI supercomputer.

  • Moore's Law on Hyperdrive: NVIDIA is pushing the boundaries of Moore's Law, achieving significant performance gains generation over generation.
  • GB200: A signed GB200 is mentioned, indicating its significance and the move towards massive production.
  • 40x Speed-up: The combination of NVIDIA's and Microsoft's innovations results in a 40x speed-up over Hopper in just two years. This compounding effect of S-curves is considered "unbelievable."

Full-Stack Innovation and System Architecture

The conversation emphasizes the importance of full-stack innovation, where advancements are made across the entire computing stack, not just the processor.

  • Evolving Architecture: The traditional model of a processor running static software is outdated. The entire stack has changed.
  • Key Components: This includes changes to the processor, NVLink (for scaling compute nodes), liquid cooling, FP4 Tensor Core architecture, and coherent connection between Grace and Blackwell.
  • NVLink: Allows scaling up to a much larger compute node.
  • Grace Blackwell: Coherently connected over a super-fast link, crucial for KV caching in large agentic models and workloads.
  • CUDA and AI Infrastructure: Combined with new CUDA algorithms and model technology on Azure's AI infrastructure, this leads to the 40x speed-up.

Fleet Management and Continuous Upgrades

Jensen emphasizes the importance of continuous upgrades to take advantage of rapid technological advancements.

  • Annual Upgrades: With technology moving at 40x per generation every two years, upgrading annually is more beneficial than waiting four years.
  • Cost Averaging: Annual upgrades allow for "performance averaging up" the entire fleet.
  • AI Factories: The complexity lies in building "giant AI factories" rather than PCs, requiring tight integration between organizations.
  • Fleet Physics: Continuously taking advantage of each year's advancements allows the fleet to provide benefits over its four- or five-year lifespan.

Software Innovation and Ecosystem Stability

The discussion highlights the compounding effects of software innovation, particularly in CUDA, and the importance of a stable ecosystem.

  • Software's Role: Software innovation adds to the hardware advancements, further enhancing performance.
  • Architecture Compatibility: NVIDIA maintains architecture compatibility across generations (Pascal, Ampere, Hopper, Blackwell), ensuring a stable and rich ecosystem.
  • Developer Investment: A large install base and rich ecosystem encourage developers to invest in enhancing their models, algorithms, and runtimes.
  • Amortization: Software investments are distributed and amortized across the entire fleet.
  • Hopper Performance Improvement: Over two years, Hopper's performance has improved by 40x through new algorithms like in-flight batching and speculative decoding.
  • Azure Container Service: Provides software gains of the fleet on the GPU fleet, allowing users to run any agent on it.
  • Runtime Optimization: Engineers from both companies are continuously optimizing runtimes, even on older architectures like Ampere (A10s, A100s).
  • Long-Term Support: NVIDIA supports and fine-tunes software for the entire life of the architecture, ensuring developer productivity and value.

Accelerated Compute and Workload Diversity

The conversation expands on the idea that GPUs are not limited to AI workloads but can accelerate a wide range of applications.

  • General-Purpose Acceleration: CUDA is both an accelerator architecture and general-purpose for heavy workloads.
  • Workload Acceleration: NVIDIA and Microsoft are collaborating to accelerate data processing (20x-50x), video transcoding, image processing, models, recommender systems, and vector search engines.
  • Fleet Utilization: As newer models move to Blackwell and future architectures, the existing fleet can be used for other workloads, ensuring full utilization.
  • Workloads per Dollar per Watt: The ultimate goal is to accelerate all workloads per dollar per watt, maximizing efficiency.

Conclusion

The discussion concludes with both Satya and Jensen emphasizing the strong partnership between Microsoft and NVIDIA and the profound impact of their combined innovations on the future of computing. They highlight the emergence of a new era of computing that was previously unimaginable and express excitement about the future.

  • Partnership and Alignment: The strong partnership and alignment between the two organizations are key to building the most advanced AI infrastructure in the world.
  • AI Transformation: AI has transformed profoundly in just two years.
  • Optimistic Outlook: Both leaders express optimism about the future, stating that "the best of times are yet to come."

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.