AI Stage - Day 1 (Google I/O 2025)

Google for DevelopersAbout 8 min readMay 21, 2025Watch original
THE SUMMARYAI-generated

Google I/O: Demis Hassabis on the Frontiers of AI & Google's AI Stack for Developers - Summary

Key Concepts

  • Frontier Models: Cutting-edge AI models pushing the limits of current capabilities.
  • AGI (Artificial General Intelligence): A theoretical AI with human-level cognitive abilities across a wide range of tasks.
  • Reasoning Paradigm: Adding a "thinking" system on top of a model to improve performance, similar to how humans think before acting.
  • Deep Think: A parallel reasoning process where multiple AI agents check each other's work.
  • Model Collapse: The potential degradation of AI model quality when trained on AI-generated content.
  • AI Watermark: An invisible marker embedded in AI-generated content to identify its origin.
  • Gemini: Google's most capable and versatile model family.
  • Gemma: Google's open-source model family.
  • JAX: A Python machine learning library used for research and building foundation models.
  • Keras: A user-friendly neural network API for applied AI.
  • XLA: A compiler for machine learning code that optimizes performance on various hardware accelerators.
  • Google AI Edge: A framework for deploying machine learning models on mobile devices and embedded systems.

Frontiers of AI (Demis Hassabis & Sergey Brin)

Improvement Potential in Frontier Models

  • Current Progress: Significant gains are being made by pushing existing AI techniques to their limits.
  • Future Breakthroughs: Achieving AGI may require one or two more major breakthroughs, with promising ideas being developed.
  • Scale vs. Algorithmic Improvement: Both scale (data and compute) and algorithmic advancements are crucial for progress. Algorithmic improvements have historically been more significant, but both are now advancing rapidly.
  • Data Centers: A significant increase in data center capacity will be needed to support the growing demand for AI models, both for training and inference.
  • Test Time Compute: Increasing the amount of computation time during inference can significantly improve model performance, especially for complex tasks.

Reasoning Paradigm and Deep Think

  • Thinking Systems: DeepMind has long been a proponent of adding "thinking" systems to AI models, as demonstrated by AlphaGo and AlphaZero.
  • Quantifiable Improvement: Turning on the "thinking" component in AlphaGo resulted in a performance leap of approximately 600 Elo points, surpassing champion level.
  • Deep Think Description: Deep Think involves parallel reasoning processes that check each other, leading to enhanced reasoning capabilities.
  • Path to AGI: Mechanisms like Deep Think are considered essential for improving reasoning and achieving AGI.
  • Creativity: Current systems lack true invention and creativity, such as solving math conjectures or proposing new physics theories.

AGI Definition and Timeline

  • Importance of AGI Definition: It's important for the field to agree on a definition of AGI.
  • AGI vs. Typical Human Intelligence: AGI should be defined as what the human brain architecture is capable of, referencing the potential of historical figures like Einstein and Mozart. Current AI systems have obvious flaws and inconsistencies compared to this standard.
  • Consistency: A true AGI should be consistently capable across a wide range of tasks, with flaws only detectable by experts after months of testing, not by individuals in minutes.
  • One Company Achieving AGI: It's likely that one company or entity will reach AGI first, but others will quickly follow due to inspiration and competition.
  • Safety and Reliability: It's crucial that the first AGI systems are built reliably and safely.
  • Emotion in AGI: Understanding emotion is necessary, but mimicking human emotions may not be required or desirable for AGI.
  • AGI Timeline: Demis estimates AGI is more likely to be achieved after 2030.

Self-Improving Systems and the Future

  • Self-Improvement Loops: The creation of self-improving systems, such as "Alpha Evolved," could accelerate AI development.
  • Pairing Techniques: Combining evolutionary programming techniques with foundation models is a promising approach.
  • Limited Game Domains: While self-improvement has been successful in limited game domains like chess and Go, its applicability to the real world remains to be seen.
  • Sergey Brin's Return to Google: The current era is a unique and scientifically exciting time for computer scientists, driving Sergey's involvement.
  • Day-to-Day Activities: Sergey focuses on the technical details of Gemini text models, pre-training, post-training, and multi-modal work.

Agents and Embodied AI

  • Visual vs. Contextual Agents: Google is interested in agents that see the world as humans do, utilizing cameras and visual input.
  • Smart Glasses: The announcement of smart glasses indicates a focus on creating assistants that understand the physical context.
  • Universal Assistant: Demis believes the universal assistant is the killer app for smart glasses.
  • Robotics: The bottleneck in robotics is software intelligence, which is now being addressed by advancements in AI models.
  • Gemini's Multi-Modal Design: Gemini was built from the beginning to be multi-modal, enabling it to understand and interact with the physical world.

Google Glass Lessons

  • Technology Gap: There was a technology gap when Google Glass was first released.
  • Consumer Electronics Supply Chain: Building consumer electronics at a reasonable price point is challenging.
  • Focus on Utility: The focus this time is on polishing the product and making it steadily available.

Video Generation and Model Quality

  • Model Collapse Concerns: There are concerns about the potential for model collapse as the internet fills with AI-generated content.
  • Data Quality Management: Google is rigorous in its data quality management and curation.
  • AI Watermarks: Google attaches invisible AI watermarks to its generative models to detect AI-generated content and combat deepfakes.
  • Synthetic Data: Eventually, video models may be good enough to be used as a source of additional training data, but care must be taken to avoid distorting the data distribution.

The Web in 10 Years and Simulation Theory

  • Web Transformation: The web will change significantly with the rise of agent-first systems.
  • Simulation Theory: Demis believes that underlying physics is information theory, suggesting a computational universe but not necessarily a straightforward simulation.
  • Anthropocentric View: Sergey argues that the concept of simulation is anthropocentric, assuming that the beings running the simulation have similar consciousness and desires to humans.

Google's AI Stack for Developers (Joana Carrasqueira & Josh Gordon)

Overview of Google's AI Stack

  • Mission: To empower every developer and organization to harness the power of AI.
  • Key Components: Robust infrastructure, state-of-the-art research, and developer tools.
  • Focus Areas: Foundation models (Gemini, Gemma), AI frameworks (JAX, Keras, PyTorch), developer tools, and infrastructure (TPU, XLA).

Foundation Models

  • Gemini Family:
    • Gemini 1.5 Pro: High context tasks, deep reasoning, excellent coding capabilities.
    • Gemini 1.5 Flash: Efficient and fast, improved reasoning, coding, and multi-modality.
    • Gemini Nano: Optimized for on-device tasks.
  • AI Studio and Gemini API:
    • New "Build" App: Generates web apps from natural language.
    • Generative Media Experience: Create and interact with creative models.
    • Built-in Usage Dashboard: Tracks API usage.
    • Native Audio TTS Support: Control emotion and style for expressive audio.
    • Enhanced Tooling: Grounding with Google Search and code execution in one API call.
    • URL Context: Provides models with content from webpages.
    • Gemini Case Support for MCP: Simplifies agentic capabilities.
  • Google AI Studio:
    • Simplest way to test deep models.
    • No Google Cloud knowledge required.
    • Free of charge.
    • Create, test, and save prompts.
    • Starter apps for inspiration.
  • Gemini Developer API:
    • Easiest way to develop with Google's foundation models.
    • Features code execution and function calling.
    • Comprehensive developer documentation and guides.
    • Gemini API cookbook with end-to-end examples.
  • GenMedia:
    • Designed to transform creative experiences across content generation.
    • Includes image, video, and audio generation models.
    • Lyria: Google's music generation model.
  • Gemma Family:
    • Gemma 2: Most advanced open model, available in 4 sizes (1, 7, 12, 27B).
    • MedGemma: Open model for multi-modal medical text and image comprehension.
    • Gemma 3n: Optimized for tablets and laptops.
    • One-Click Deployment: Deploy Gemma models from AI Studio to Cloud Run.

AI Frameworks

  • Keras:
    • Easiest way to apply AI in practice.
    • Fine-tune models with a 2-column CSV file.
    • 5 key lines of code to import and fine-tune a Gemma model.
  • JAX:
    • Python machine learning library with NumPy API.
    • Scales easily to tens of thousands of accelerators.
    • Used to build Gemini and Gemma.
    • Ecosystem of libraries for optimizers, checkpoints, and network implementation.
    • MaxText and MaxDiffusion: Reference implementations of large language models and diffusion models.
    • Marin: Fully open model released by Stanford University, built with JAX and TPUs.
    • Tunics: New library for post-training algorithms in JAX.
  • PyTorch:
    • Works with XLA for optimized performance.
    • vLLM: Super-popular inference engine with TPU support for PyTorch.

Infrastructure

  • XLA (Accelerated Linear Algebra):
    • Compiler for machine learning code.
    • Optimizes code for GPU and TPU.
    • Portable and hardware-agnostic.
  • LLM D: Partnership between Red Hat, Nvidia, and Google for distributed serving of large language models.

Google AI Edge

  • Framework for deploying machine learning on mobile devices and embedded systems.
  • Benefits:
    • Low latency.
    • Privacy.
    • Offline functionality.
    • Cost savings.
  • Features:
    • Support for Gemma models.
    • Community of Hugging Face users.
    • AI Edge Portal: Testing service for verifying model performance on real devices.

Future Domains

  • Alpha Evolve: Gemini coding agent for designing advanced algorithms.
  • AI Co-Scientist: Accelerates scientific discovery and drug development.
  • Gemini Robotic Models: Advanced vision, language, action models for robotics.

Synthesis/Conclusion

The Google I/O session highlighted Google's commitment to pushing the boundaries of AI and making it accessible to developers. Key takeaways include the importance of both scale and algorithmic improvements in AI development, the potential of reasoning paradigms and self-improving systems, and the focus on building safe and reliable AGI systems. Google's AI Stack provides a comprehensive set of tools and frameworks for developers to build innovative applications across various domains, from mobile devices to cloud infrastructure. The emphasis on open-source models like Gemma and community collaboration further underscores Google's dedication to fostering a vibrant and inclusive AI ecosystem.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.