How Nvidia GPUs Compare To Google’s And Amazon’s AI Chips

By CNBC

Share:

Key Concepts

  • GPUs (Graphics Processing Units): Originally designed for rendering graphics, their parallel processing capabilities make them ideal for AI training and inference.
  • ASICs (Application-Specific Integrated Circuits): Custom-designed chips optimized for a single purpose, offering high efficiency and speed for specific AI workloads.
  • FPGAs (Field-Programmable Gate Arrays): Reconfigurable chips that can be adapted for various applications, including AI, offering flexibility but lower raw performance than ASICs.
  • Edge AI: AI processing that occurs on local devices (smartphones, cars, etc.) rather than in the cloud, often utilizing NPUs.
  • NPUs (Neural Processing Units): Dedicated AI accelerators integrated into SoCs (Systems on a Chip) for edge AI devices.
  • Tensor Processing Units (TPUs): Google's custom ASICs designed for AI acceleration.
  • CUDA: Nvidia's proprietary software platform that optimizes GPU performance.
  • TSMC (Taiwan Semiconductor Manufacturing Company): The primary manufacturer of advanced AI chips for most major tech companies.
  • AI Training: The process of teaching an AI model by exposing it to large datasets.
  • AI Inference: The process of using a trained AI model to make predictions or decisions on new data.
  • SoC (System on a Chip): A single integrated circuit that combines multiple components, such as a CPU, GPU, and NPU, often found in mobile devices.

Nvidia's Dominance and the Rise of AI Chips

Nvidia has transitioned from a gaming company to a central player in generative AI, with its Blackwell GPUs powering AI workloads globally. The company's valuation has surged, reaching $5 trillion briefly, driven by the demand for its GPUs. Nvidia's Blackwell GPUs, with 6 million shipped in the last year, are designed to connect up to 72 GPUs to function as a single unit for advanced AI tasks.

GPUs: The Workhorse of AI

GPUs, like those from Nvidia and its competitor AMD, are the primary workhorses for AI. Their architecture, optimized for parallel programming, is well-suited for AI tasks that require simultaneous calculations, such as rendering images or training neural networks. The "big bang moment" for AI, AlexNet in 2012, demonstrated the effectiveness of GPUs for image recognition.

  • Technical Detail: GPUs possess thousands of smaller cores focused on parallel math, particularly matrix multiplication, which is crucial for processing tensors (multidimensional data structures). This contrasts with CPUs, which have fewer, more powerful cores for sequential tasks.
  • Key Argument: The parallel processing capability of GPUs, originally for graphics, is highly effective for training neural networks, where computers learn from data.
  • Example: Researchers in 2012 "hacked" GPUs to expose their parallel computation capabilities for deep learning.

AI Workloads: Training vs. Inference

AI workloads typically involve two main phases: training and inference.

  • Training: Teaching an AI model to recognize patterns in vast datasets. This is compute-intensive and a primary use case for GPUs.
  • Inference: Using a trained AI model to make decisions or predictions on new information. This is how AI is experienced in everyday applications like smartphone apps or voice assistants. Inference can be performed on less powerful, more specialized chips.

Nvidia's Business Model and Market Position

Nvidia sells GPUs directly to AI companies (e.g., a deal for 4 million GPUs to OpenAI) and governments, as well as to cloud providers like Amazon, Microsoft, and Google, who then rent them out. Nvidia also offers its own GPU rental program. The demand for AI systems is so high that a 72-GPU Blackwell server rack can cost around $3 million, with Nvidia shipping 1,000 per week. Nvidia aims to sell complete systems for greater efficiency.

  • Data: Nvidia is shipping 1,000 Blackwell server racks per week.
  • Quote: "And for Nvidia, they're looking to sell the entire system, not just the chip. Because when you think of that system level, you get more efficiencies in terms of speed and power and performance."

AMD: Nvidia's GPU Competitor

AMD is Nvidia's main competitor in the GPU market, with its Instinct GPU line seeing significant adoption, including commitments from OpenAI and Oracle. A key differentiator for AMD is its largely open-source software ecosystem, while Nvidia relies on its proprietary CUDA platform. Nvidia's next-generation Rubin GPU is slated for full production next year.

The Rise of Custom ASICs

As AI models mature and the need for inference grows, custom ASICs are gaining significant traction. These chips are designed by major hyperscalers like Google, Amazon, Meta, and Microsoft, as well as companies like Broadcom.

ASIC Characteristics and Advantages

  • Definition: Application-Specific Integrated Circuits (ASICs) are custom-designed chips built for a single purpose.
  • Advantages: ASICs are smaller, cheaper, and more power-efficient for their specific tasks compared to general-purpose GPUs. They can be significantly more efficient in running targeted workloads.
  • Trade-off: The specialization of ASICs means they are hard-wired and cannot be changed once manufactured, leading to a trade-off in flexibility.
  • Market Growth: The ASIC market is projected to grow even faster than the GPU market in the coming years.

The Cost and Strategic Importance of ASICs

Developing a custom ASIC is extremely expensive, costing tens to hundreds of millions of dollars. This makes them inaccessible for startups, which often rely on GPUs. However, for large cloud providers, ASICs offer long-term cost savings, power efficiency, and reduced reliance on external chip suppliers like Nvidia. They also provide greater control over their AI workloads.

  • Cost: Custom ASICs cost a minimum of tens, and often hundreds of millions of dollars.
  • Strategic Goal: Hyperscalers aim to reduce the cost of AI and gain more control over their workloads.

Key Players in the ASIC Space

  • Google: Pioneered custom AI ASICs with its Tensor Processing Unit (TPU) in 2015. The TPU was instrumental in the development of the transformer architecture, which powers most modern AI. Google's latest generation is Ironwood, with a deal to train Anthropic's Claude LLM on up to a million TPUs.
    • Technical Detail: Google's chip lab in 2024 showcased Trillium chips, with two connected to a host machine and others linked to form a large supercomputer.
    • Speculation: There's speculation that Google may offer broader access to TPUs in the future.
  • Amazon Web Services (AWS): Developed its own AI chips after acquiring Annapurna Labs. AWS announced Inferentia in 2018 and launched Trainium in 2022, with a third generation approaching.
    • Comparison: Trainium is described as a cluster of workshops with flexible tensor engines, while Google's TPU is like a factory conveyor belt with a rigid grid for matrix math.
    • Performance: Trainium offers 30-40% better price performance compared to other hardware vendors on AWS.
    • Real-world Application: Anthropic is training its models on half-a-million Trainium2 chips in Amazon's data centers, notably without Nvidia GPUs in that specific setup.
  • Meta: Launched its training and inference accelerator in 2023.
  • OpenAI: Has a significant new deal with Broadcom to build custom ASICs starting in 2026.
  • Microsoft: Is developing its own Maia chips for Azure data centers, though its next chip is facing delays.
  • Intel: Offers its Gaudi line of custom ASICs.
  • Tesla: Has announced its own ASIC.
  • Qualcomm: Is entering the data center chip market with its AI200.
  • Startups: Companies like Cerebras (large wafer-scale AI chips) and Grok (inference-focused language processing units) are also developing custom AI chips.

The Role of Chip Design Companies

Companies like Broadcom and Marvell act as backend partners, providing intellectual property (IP), know-how, and networking expertise, allowing hyperscalers to avoid building full silicon teams. Broadcom is a major beneficiary of the AI boom, having helped build Google's TPUs and Meta's accelerators, and is now partnering with OpenAI.

  • Market Share: Broadcom is estimated to hold 70-80% of the ASIC design market.
  • Market Growth: This market is expected to accelerate at a mid-double-digit CAGR over the next five years.

Edge AI and FPGAs

Edge AI Chips (NPUs)

Another significant category of AI chips is for running AI on edge devices (e.g., smartphones, cars, smart home devices) rather than in the cloud. This allows for local AI processing, enhancing privacy and responsiveness.

  • Definition: Neural Processing Units (NPUs) are dedicated AI accelerators integrated into a device's primary chip (SoC).
  • Advantages: NPUs are integrated, take up less silicon, and are therefore less expensive than data center chips. They enable AI processing without constant cloud communication.
  • Key Players: Qualcomm, Intel, and AMD are major NPU manufacturers for PCs. Apple's M-series chips for MacBooks include a "neural engine," and its iPhone A-series chips also have dedicated neural accelerators. Android phones with Qualcomm Snapdragon and Samsung Galaxy phones feature NPUs.
  • Applications: NPUs power AI in cars, robots, cameras, and smart home devices.
  • Future Trend: While data center AI currently receives most attention, edge AI is expected to grow significantly as AI is deployed in more devices.

FPGAs (Field-Programmable Gate Arrays)

FPGAs are reconfigurable chips that can be reprogrammed after manufacturing for various applications, including AI.

  • Flexibility: More flexible than ASICs or NPUs.
  • Trade-offs: Lower raw performance and energy efficiency for AI workloads compared to ASICs.
  • Use Case: Chosen when designing a custom ASIC is not feasible, but ASICs are cheaper for large-scale deployments.
  • Market: AMD became the largest FPGA maker after acquiring Xilinx for $49 billion, and Intel is second after acquiring Altera for $16.7 billion.

Chip Manufacturing and Geopolitics

The manufacturing of advanced AI chips is heavily concentrated with Taiwan Semiconductor Manufacturing Company (TSMC). Companies like Nvidia, Google, and Amazon rely on TSMC for production.

  • Geopolitical Significance: The reliance on Taiwan for chip manufacturing presents geopolitical challenges.
  • US Investment: The US is investing in domestic chip fabrication, with TSMC opening a plant in Arizona. Apple has committed to some production there, though its latest iPhone chip uses TSMC's 3nm node, currently only available in Taiwan. Nvidia's Blackwell GPUs are made on TSMC's 4nm node, with production now in Arizona. Intel is also reviving its foundry business with advanced chip manufacturing in Arizona.
  • "Silicon Back to Silicon Valley": The AI boom is driving a resurgence in chip manufacturing in the US.

Chinese AI Chip Development

Major Chinese players like Huawei, ByteDance, and Alibaba are developing custom ASICs. However, they face limitations due to export controls on advanced equipment and AI chips.

Energy and Global Competition

A critical factor for the massive AI data center build-out is securing sufficient power. The US faces energy risks in this regard, while China has been more proactive.

  • Quote: "If the US wants to continue to lead in AI, we continue to be fraught with energy risk. China has done that much better than us, for instance."
  • Quote: "And while we do have the best chips in the world, and I believe so by multiple generations, I believe that our need to build out energy is critical."

The Race for AI Supremacy

Despite Nvidia's current lead as the world's most valuable company, the race to develop AI chips is intensifying. While dethroning Nvidia will be difficult due to its established developer ecosystem, the sheer size of the AI market ensures continued new entrants.

  • Quote: "Although dethroning Nvidia won't come easily. They have that position because they've earned it and they've spent the years building it, and they've won that developer ecosystem. But that market's going to get so big that we're going to continue to see new entrants."

Conclusion

The AI chip landscape is rapidly evolving, with GPUs like Nvidia's remaining dominant for high-performance training, while custom ASICs are emerging as a more efficient solution for specific inference tasks, driven by hyperscalers seeking cost reduction and control. Edge AI, powered by NPUs, is poised for significant growth, bringing AI processing closer to the user. The global race for AI leadership is also intertwined with manufacturing capabilities, geopolitical considerations, and the critical need for energy infrastructure. While Nvidia holds a strong position, the immense market potential guarantees continued innovation and competition across all categories of AI chips.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video