What's Next for AI Infrastructure with Amin Vahdat | AI Basics with Google Cloud

By This Week in Startups

Share:

Key Concepts

TPUs (Tensor Processing Units), GPUs, AI infrastructure scaling, deep research, inference, training, model serving, unit cost of intelligence, engineering productivity, AI agents, code generation, cloud computing, exponential growth, AI-driven automation, supervised associate agents.

Infrastructure Scaling and Deep Research

Amin Vahdat, VP and GM of Machine Learning and Cloud AI at Google Cloud, discusses the scaling of AI infrastructure, particularly focusing on TPUs and GPUs. He uses the example of a complex music recommendation query using Google Gemini's deep research capabilities to illustrate the massive compute power involved.

  • Scale of Compute: A "normal" web search involves 1,000-10,000 standard computers. Deep research queries utilize TPUs, custom accelerators that pack the computing power of approximately 100 standard servers into a single chip. A single query might utilize 256 TPUs, equating to tens of thousands of server equivalents.
  • Iterative Process: Deep research involves multiple sub-queries (potentially 10-100) that are run and composed in real-time to provide a comprehensive answer.
  • Data Movement: The discussion emphasizes the exponential increase in bandwidth and storage capabilities compared to the early days of the internet, where even email signatures impacted bandwidth.

The Age of Inference

The conversation shifts to the increasing importance of inference, the process of serving models and processing queries in real-time.

  • Shift from Training to Inference: The focus has moved from training models to efficiently serving them.
  • Rapid Efficiency Gains: Google is achieving efficiency improvements at a rapid pace, with potential 2x speed increases every 3 months, leading to exponential compounding.
  • Hardware and Software Synergies: Improvements are driven by both software optimizations and advancements in hardware.

Startup Applications and Productivity

The discussion explores how startups are leveraging the massive AI infrastructure.

  • Productivity Focus: The hottest areas for AI application are currently around engineering and employee productivity, particularly in code generation and bug finding.
  • Developer Productivity: AI is seen as a tool to multiply the productivity of developers rather than replace them entirely.
  • Transformative Engineering Talent: The constraint for both large companies and startups is transformative engineering talent.

Pricing and Cost Reduction

The conversation addresses the decreasing costs of AI compute.

  • Unit Cost of Intelligence: Google is focused on driving down the "unit cost of intelligence," which is the cost normalized to quality.
  • Significant Cost Reductions: Cost reductions are happening at a much faster rate than traditional storage, potentially a factor of three reduction (300%) in a year.
  • Infrastructure Availability: Infrastructure availability is no longer the primary bottleneck for startups.

Historical Perspective and Future Possibilities

The discussion draws parallels to the early days of the internet to illustrate the transformative potential of current AI advancements.

  • Gmail, Flickr, and YouTube Examples: These services were initially considered non-viable due to storage and bandwidth costs.
  • Unlimited Models: Entrepreneurs who embraced "unlimited" models disrupted the market.
  • Thinking Exponentially: Entrepreneurs should challenge assumptions of impossibility by dividing cost equations by factors of 10 or 100 to account for future advancements.

TPUs and Google's Approach

The conversation delves into the history and development of TPUs at Google.

  • Voice Recognition Use Case: The initial motivation for TPUs was the projected compute demand for voice recognition in 2013.
  • Matrix Multiplication Efficiency: TPUs were designed to perform matrix multiplications, a core operation in machine learning, 100 times more efficiently than CPUs.
  • Generational Improvements: Google has developed seven generations of TPUs, with the current generation offering a 10x performance improvement over the previous one.

Unexpected Applications and the Rise of Agents

The discussion explores unexpected applications of AI and the emerging role of AI agents.

  • Exciting Applications: The speaker sees exciting applications in how things are progressing with agents.
  • AI Agents: AI agents are now able to invoke code, interact with other agents, and take actions based on generated answers.
  • Automation and Delegation: The speaker emphasizes the importance of automating, deprecating, and delegating repetitive tasks.

Supervised Associate Agents

The conversation concludes with a specific example of an AI agent being developed for venture capital deal analysis.

  • Deal Analysis Agent: The agent will clean and categorize startup data, check other sources, and build a dossier.
  • Brutal Candor: The AI will provide brutally candid feedback, identifying missed opportunities and potential blind spots.
  • Reinforcement Learning: The AI can learn from past decisions and provide insights to improve future decision-making.

Synthesis/Conclusion

The interview highlights the rapid advancements in AI infrastructure, particularly driven by TPUs and the shift towards efficient inference. The decreasing costs and increasing capabilities are empowering startups to build innovative applications, especially in areas like engineering productivity and AI-driven automation. The emergence of AI agents capable of taking actions and providing candid feedback represents a significant step towards more sophisticated and impactful AI systems. The key takeaway is that the constraints of the past are rapidly disappearing, and the primary limitation is now human imagination and the ability to effectively utilize the available infrastructure.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video