OpenAI Targets Custom Silicon in Broadcom Deal

By Bloomberg Technology

Share:

Key Concepts

  • Golden Touches/Sweetener: Favorable terms or special deals offered to a company.
  • AI Chips: Specialized processors designed to accelerate artificial intelligence workloads.
  • Google TPU (Tensor Processing Unit): Google's proprietary custom-designed AI accelerator hardware, optimized for machine learning.
  • Merchant Silicon: Off-the-shelf, commercially available semiconductor chips (e.g., GPUs from Nvidia or AMD).
  • Custom Silicon: Semiconductor chips designed specifically for a particular application, customer, or workload, offering tailored performance and efficiency.
  • Hyperscalers: Large cloud computing providers (e.g., Google, Amazon, Microsoft) that operate at massive scale.
  • Inference: The process of using a trained AI model to make predictions or decisions on new data, as opposed to training the model.
  • Large Language Model (LLM): An AI model trained on a vast amount of text data, capable of understanding, generating, and processing human language.
  • Performance per Watt: A metric measuring the computational work performed per unit of power consumed, indicating energy efficiency.
  • Latency: The delay between a request and a response in a system, crucial for real-time AI applications.
  • Gigawatt: A unit of power, here used metaphorically to represent the scale of a large data center buildout, often associated with significant energy consumption and cost.

OpenAI's Strategic Partnership with Broadcom

OpenAI has secured "golden touches" or a "sweetener" deal with Broadcom, a unique arrangement as such favorable terms are typically not extended by major chip manufacturers like Nvidia or AMD to other publicly traded companies. This partnership signifies OpenAI's strategic move to secure custom silicon, mirroring Google's highly successful approach with its Tensor Processing Units (TPUs).

The Google TPU Model: A Blueprint for Success

The underlying model for OpenAI's strategy is the Google TPU model. Broadcom's AI chip business currently boasts an almost $20 billion run rate, with more than half of that revenue derived from Google's TPUs. Google has developed its TPUs at a very quick pace, now in their seventh version of chips, demonstrating a successful long-term strategy for custom silicon. This success has positioned Google with the lowest infrastructure cost among all hyperscalers. OpenAI aims to replicate this "ramp up" success.

Cost Reduction and Efficiency Gains

A primary driver for OpenAI's collaboration with Broadcom is significant cost reduction. By leveraging Broadcom's custom chips, OpenAI anticipates reducing costs by 30% to 40% per gigawatt of data center buildout. A single gigawatt data center typically costs $40 billion to $50 billion. The cost of chips constitutes the highest component in such a buildout, accounting for 60% to 70% of the total data center cost. Broadcom's custom silicon can substantially lower this chip cost, making the overall infrastructure significantly cheaper. This strategy combines the benefits of "merchant silicon" (for general availability) with "custom silicon" (for specialized optimization), similar to Google's diversified approach.

Competitive Landscape and Custom Silicon Advantages

The transcript highlights the competitive landscape for custom AI silicon:

  • Google + Broadcom (TPU): The benchmark for success, achieving low infrastructure costs and high performance.
  • Amazon + Marvell (Cranium): Amazon's custom chip effort with Marvell has not achieved the same level of success as Google's TPUs.
  • Microsoft + Marvell: Similarly, Microsoft's efforts with Marvell have not matched Google's success.
  • Nvidia: While a major player, Nvidia remains the highest cost chip provider, even with its investments.
  • AMD: AMD is likely to cut costs but is not expected to match the performance per watt of custom solutions.

Given these comparisons, OpenAI's choice of Broadcom is seen as a natural fit, leveraging Broadcom's proven ability to deliver custom specifications at scale, as demonstrated with Google.

Technical Rationale for Custom Silicon

The need for custom silicon is driven by specific technical requirements for running large language models (LLMs) efficiently, particularly for inference. Key optimization goals include:

  • Minimum Latency: Reducing the delay in processing, critical for responsive AI applications.
  • Maximum Performance per Watt: Achieving the highest computational output for the lowest power consumption, addressing power as a significant constraint.

The discussion mentions exploring concepts like "tiny recursive models" to run large models more efficiently. Custom silicon, as exemplified by Google's TPUs, allows for tailored optimization to achieve these goals. Google's custom silicon, for instance, enables the best performance for running YouTube videos, a level of optimization that "no other merchant silicon can give." OpenAI is pursuing this same level of tailored performance for its LLMs.

Conclusion: Replicating Google's Efficiency for AI Scale

OpenAI's partnership with Broadcom for custom silicon is a strategic and financially driven decision to emulate Google's success with TPUs. By securing custom-designed chips, OpenAI aims to drastically reduce its AI infrastructure costs (by 30-40% per gigawatt) and significantly improve the performance per watt of its large language models, especially for inference workloads. This move is critical for scaling AI operations efficiently, minimizing latency, and managing the substantial cost of data center components, where chips represent the largest expenditure. The ability of Broadcom to provide custom specifications and scale production, as proven with Google, makes it an ideal partner for OpenAI's ambitious goals.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video