Gemini co-leads on project origins and what's next

Google for DevelopersAbout 4 min readMay 30, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini: A unified, multimodal, general-purpose AI model series developed by Google DeepMind.
  • Multimodality: The ability of a model to process and understand various data types (text, images, audio, video, code, and scientific data).
  • Agentic AI: AI systems capable of autonomous decision-making, tool use, and executing complex, multi-step tasks.
  • Distillation: A technique where a smaller, more efficient model (student) learns from a larger, more powerful model (teacher) to retain high performance with fewer parameters.
  • World Models: AI systems that understand the physical dynamics and physics of the world, allowing them to simulate future states.
  • Pathways Project: A foundational Google initiative aimed at creating a single, sparse, multimodal model capable of performing many tasks.
  • One Box Philosophy: The concept of a single, unified interface (like a search box) that provides diverse services, now powered by a single, powerful AI backend.

1. The Origin and Evolution of Gemini

The Gemini project was born out of a strategic decision to consolidate fragmented AI efforts across Google (including DeepMind and the Pathways project). Jeff Dean noted that fragmenting compute and research talent was inefficient. By merging these teams, Google created a "single, powerful model" that could serve as the core engine for Google’s intelligence. The name "Gemini" (the twins) reflects this unification of previously separate research streams.

2. The "Gemini 3.5" Era and Technical Strategy

  • Focus: The 3.5 series, specifically the "Flash" model, emphasizes coding capabilities and agentic experiences.
  • Product-Driven Research: The speakers argue that building AI in a "black box" (focusing only on benchmarks) is insufficient. Real-world product integration is essential to understand user needs and identify where models fall short.
  • Efficiency: A major technical achievement is the ability to pack high-level intelligence into smaller models (Flash) through advanced distillation, allowing for high performance with lower latency.

3. World Models and Multimodality

The team discussed the transition from simple text-to-video generation to "true world models."

  • Definition: A world model understands physical dynamics and can simulate future states.
  • Omni: Gemini Omni represents a shift toward training jointly across modalities, enabling the model to understand the physical world alongside text, which provides high-level conceptual context.
  • Beyond Human Modalities: The goal is to expand beyond text/audio/video to include scientific data, such as genomic sequences, chemical structures, and robotic sensor data (LiDAR).

4. Organizational History and Collaboration

The discussion highlighted the deep-rooted history of the team:

  • Mentorship: Jeff Dean and others played pivotal roles in recruiting and mentoring key researchers like Oriel and Oriol, fostering a culture of shared knowledge.
  • The "Code Review" Culture: The team emphasized that early collaboration was built on hands-on code reviews and shared engineering challenges, which helped bridge the gap between academic research and production-scale engineering.
  • Overcoming Geography: Despite the 8-hour time difference between London and the US, the team successfully integrated by focusing on shared goals and data-driven experimentation.

5. Challenges and Future Outlook

  • Evaluation: The team identified evaluation as a significant, underappreciated challenge. Moving from static "table of numbers" benchmarks to real-world user feedback is critical for measuring true generalization.
  • Continual Learning: A key area for future improvement is enabling models to learn from experience without needing constant weight updates.
  • Efficiency: There is a massive gap between human learning efficiency and LLM training. The team aims to extract more information from every token/example to improve efficiency.
  • Predictions for 2027:
    • Self-Learning: Models will likely begin to assist in their own research and development.
    • Long-Running Agents: The ability for models to run autonomously for extended periods (e.g., 30 days) to complete complex tasks.
    • Tool Latency: As models become faster, the bottleneck will shift to the latency of the tools they interact with, which are currently designed for human speeds.

6. Notable Quotes

  • Jeff Dean: "We are fragmenting our efforts and fragmenting our compute and if we're going to build an incredibly powerful model, we need to all come together and work on building a single model."
  • Oriol Vinyals: "You don't want to build intelligence in a black box. You want it to be useful... those two [research and product] hand in hand define what the frontier means."
  • Jeff Dean (on distillation): "It's like squeezing the lemon. You squeeze the lemon, the juice comes out, it's the good bits, you put it in a glass, which is your small model."

Synthesis

The Gemini project represents a shift from fragmented, academic-style AI research to a unified, product-integrated engineering powerhouse. By focusing on multimodality, agentic capabilities, and efficient distillation, the team is moving toward "world models" that can simulate reality. The future of AI at Google, according to the team, lies in self-improving systems that can operate autonomously over long durations, ultimately transforming how humans interact with information and the physical world.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.