Key Concepts
- Gemini 3: The latest iteration of Google's AI model, with positive reception and advancements across various dimensions.
- AGI (Artificial General Intelligence): The ultimate goal of building intelligence that can understand, learn, and apply knowledge across a wide range of tasks at a human-like level.
- Co-building AGI: The philosophy that AGI is not solely developed in isolation but is built collaboratively with users and customers through product integration and feedback.
- Innovation and Ideas: The driving force behind AI progress, fueled by real-world applications and user signals.
- Benchmarks: Tools for guiding model development and measuring progress, but not the sole definition of advancement.
- Instruction Following: A critical capability for AI models to accurately understand and execute user requests.
- Internationalization: Ensuring AI models are effective and accessible across diverse languages and regions.
- Function Calls/Tool Calls/Agentic Actions: Enabling models to interact with external tools and services to enhance their capabilities.
- Code Generation: A key area for building applications and bringing digital ideas to life.
- Product Scaffolding: The importance of products like Anti-Gravity and AI Studio in providing platforms for model improvement and user feedback.
- Engineering Mindset: Applying robust engineering principles to AI development, focusing on safety, security, and reliability from the ground up.
- Collaborative Effort: The realization that building advanced AI models requires a massive, global, and cross-functional team effort.
- Multimodality: The ability of AI models to process and understand information from various modalities, including text, images, and audio.
- Generative Models: AI models capable of creating new content, historically focused on images and now expanding to other modalities.
- Unified Model Checkpoints: The aspiration to have a single, cohesive model architecture that can handle diverse tasks and modalities.
- Scientific Exploration vs. Scaling: The ongoing challenge of balancing fundamental research and scientific discovery with the practical scaling and deployment of AI models.
- Underdog Story: The narrative of Google's journey in AI, evolving from a position of catching up to a leadership role through innovation and strategic investment.
Gemini 3: A New Frontier in AI
The discussion centers around the recent launch of Gemini 3, highlighting its overwhelmingly positive reception and the significant progress made by DeepMind and Google AI. The sentiment is one of excitement and validation, with the belief that this advancement is a crucial step towards building Artificial General Intelligence (AGI).
Key Points on Gemini 3 and Progress
- Positive Reception: The model's launch has been met with "super positive" reception, with users finding the model's capabilities and the aspects that the development team found interesting to be equally compelling.
- Continuous Progress: A key theme is the observation that AI progress is not slowing down. The launch of Gemini 3 follows the advancements seen with Gemini 2.5, indicating a sustained pace of innovation.
- Pushing the Frontier: Gemini 3 is described as having "pushed the frontier on a bunch of dimensions," mirroring the sentiment from the Gemini 2.5 launch.
- Research and Innovation: The progress is attributed to continuous innovation and new ideas across all areas of AI development, from data pre-training to post-training. The real-world application of models generates more ideas and increases the "surface area" for learning.
- Challenging Problems: As AI capabilities grow, the problems become more complex and varied, which is seen as a positive driver for further development towards AGI.
Benchmarks and Defining Progress
The conversation delves into the role and limitations of benchmarks in measuring AI progress.
- Benchmarks as Guides: Benchmarks are essential for guiding model development, but they are often defined at a specific point in time and can become saturated as technology advances.
- Evolving Frontiers: As models master existing benchmarks, new frontiers and benchmarks need to be defined to accurately reflect the state-of-the-art.
- Examples of Benchmark Advancement:
- HLE (Human Language Evaluation): Models have progressed from performing poorly (1-2%) to achieving over 40% with "deep think."
- RKGI2: Similarly, models have moved from struggling to perform any tasks to achieving over 40%.
- Static Benchmarks: While some static benchmarks like GPQA (Graduate-Level Google-Proof Questions) continue to be relevant, progress is incremental (e.g., "eking out 1%"). These benchmarks still test difficult problems, but models are getting closer to solving them.
- Defining Progress Beyond Benchmarks: The most important measure of progress is seen in the real-world application of models by scientists, students, lawyers, engineers, and the general public for various tasks like writing, creative endeavors, and email composition. The ability to deliver value across a spectrum of difficulty and domains is paramount.
Strategic Areas for Improvement ("Hill Climbing")
The discussion highlights specific areas where Gemini models, particularly the Pro version, are being optimized.
- Instruction Following: Ensuring models accurately understand and adhere to user requests, rather than providing answers they deem appropriate.
- Internationalization: Expanding model capabilities to support a wider range of languages, addressing historical limitations and aiming for global accessibility. Gemini 3 Pro has shown significant improvements in languages where Google historically struggled.
- Technical Domains:
- Function Calls and Tool Calls: Enhancing models' ability to seamlessly use existing tools and functions, and even write their own, acting as a "multiplier of intelligence."
- Agentic Actions: Developing models that can act autonomously and perform complex tasks.
- Code Generation: A critical area for building applications and enabling users to "bring anything to life" in the digital world. This is seen as a foundational element for integration with various aspects of users' lives.
- "Wipe Coding" Example: This illustrates how code generation empowers creative individuals to translate ideas into functional applications quickly, bridging the gap between creativity and productivity. This enables more people to become "builders."
The Role of Product Scaffolding and User Feedback
The integration of AI models into products is crucial for their development and refinement.
- Anti-Gravity and AI Studio: These platforms are highlighted as important "product scaffolding" that allows for "hill climbing" on model quality.
- Direct User Learning: Products like Anti-Gravity, Gemini App, AI Studio, and AI Overviews in Search provide direct integration with end-users, enabling developers to learn from them and understand where models need improvement.
- Critical Feedback: The feedback from these product integrations has been "instrumental" in the development process, providing real-world use case signals that complement benchmark data.
- Bridging the Gap: This integration helps bridge the gap between scientific benchmarks and real-world utility, ensuring models are useful and impactful.
The Chief AI Architect Role and Cross-Google Integration
Cory Cavachulu's new role as Chief AI Architect of Google emphasizes the integration of DeepMind's technology across all Google products.
- Enabling Products: The focus is on making the best available technology accessible to Google's product teams, rather than developing products themselves.
- Defining the New World: This new AI technology is redefining user expectations and how products should behave.
- Co-building AGI with Customers: The philosophy of co-building AGI extends to working with other product areas and the broader ecosystem.
- Engineering Mindset: A strong emphasis is placed on an engineering mindset for building robust, safe, and reliable AI systems, integrating safety and security from the initial stages of development.
- Team Google Effort: The development of Gemini models is a massive, global, and collaborative effort involving teams across all of Google, not just DeepMind. This includes product teams working in parallel with model development.
The Future of Gemini and Beyond
Looking ahead, the discussion touches upon future aspirations and challenges.
- Continuous Improvement: Despite Gemini 3's success, areas like writing and coding still have room for improvement, particularly in agentic actions and coding.
- Evolution of Agentic Actions and Tool Use: The progress in agentic tool use is seen as a significant growth area, with a focus on improving capabilities beyond current state-of-the-art.
- Historical Focus: The initial focus on multimodality for Gemini 1.0 is acknowledged, with a gradual shift towards agentic infrastructure in 2.0. The rate of progress in agentic tool use is strong, but it's a continuous effort.
- Shift from Research to Engineering: The journey of Gemini represents a shift from a pure research environment to an engineering mindset, with a strong connection to products and users. This involves rapid iteration and monthly updates.
- Multimodality and Generative Models:
- Historical Context: Generative models were initially focused on images due to better inspectability and the drive to understand the world and physics.
- Text as a Catalyst: Text emerged as a domain for rapid progress, but image and video models are now returning to prominence.
- Convergence of Architectures: Architectures for different modalities are merging, leading to more efficient and integrated models.
- Nano Banana: This model is presented as an early example of this convergence, allowing for image iteration and interaction with text models, combining world understanding from both perspectives.
- Nano Banana Pro: This advanced image generation model, built on Gemini 3 Pro, demonstrates enhanced performance in nuanced use cases like text rendering and world understanding.
- Unified Model Checkpoints: The goal is to move towards unified Gemini model checkpoints, where different modalities are integrated into a single model architecture. This is a challenging but achievable goal requiring significant innovation.
- Output Space and Learning Signal: The output space is critical for learning signals. While text and code have been strong drivers, achieving high-quality and conceptually coherent image generation is a complex task.
DeepMind's Legacy and the Journey to AGI
The conversation reflects on DeepMind's historical contributions and its evolution.
- Pioneering Deep Learning: Cory Cavachulu was the first deep learning researcher at DeepMind, starting 13 years ago when the technology was not widely embraced.
- Visionary Beginnings: DeepMind was visionary in its focus on building intelligence with deep learning at its core.
- Learning from Past Projects: DeepMind has learned valuable lessons in organizing around goals and missions from projects like DQN, AlphaGo, AlphaZero, and AlphaFold.
- Engineering Mindset Integration: The last few years have seen the integration of an engineering mindset, with a focus on developing and exploring a mainline of models.
- Deep Think Models: These models, used in competitions like IMO, exemplify the approach of exploring and evolving existing models for specific, challenging targets, making advanced capabilities accessible to a wider audience.
- Large-Scale Collaboration: The scale of contributions has grown significantly, with Gemini 3 involving thousands of contributors, reflecting the complexity and breadth of modern AI development.
- Full-Stack Approach: Google's full-stack approach, from data centers to chips and networking, is a significant advantage, enabling the design of models and hardware in tandem.
- Balancing Exploration and Scaling: A critical challenge is balancing pure scientific exploration with the scaling of models like Gemini. The biggest risk for Gemini is "running out of innovation."
- Innovation as the Engine: Innovation, at various scales and directions, is the key to achieving the goal of building intelligence. This includes exploring new architectures and ideas within the Gemini project and broader DeepMind/Google Research efforts.
- Gemini as a Goal, Not an Architecture: Gemini is defined by the goal of achieving intelligence, not a specific architecture, allowing for flexibility and evolution.
Culture and Collaboration
The importance of culture and human connection in AI development is also discussed.
- Humanity and Warmth: The collaborative spirit and the "warmth of humanity" are seen as integral to the process of launching AI models and driving innovation.
- Deep Scientific Roots and Kindness: DeepMind's culture, influenced by Demis Hassabis, emphasizes deep scientific roots combined with kindness and friendliness.
- Trust and Opportunity: The belief in the team, trusting individuals, and providing opportunities are key to fostering a collaborative environment.
- Caring About Impact: The environment fosters a sense of care for solving challenging technical and scientific problems that have a real-world impact.
- Humility and Self-Questioning: Approaching the goal of building intelligence with humility and a willingness to question oneself is crucial.
- Teamwork and Support: Despite exhaustion, the team's ability to come together, support each other, and tackle hard problems makes the process enjoyable and effective.
- Potential of Technology: Acknowledging the potential of technology while recognizing that future architectures may differ from current LLMs. Pushing for new exploration in multiple directions is seen as the right approach.
- Capabilities Speak for Themselves: The focus should be on the capabilities and demonstrations of AI in the real world, rather than debates about what is "right" or "wrong."
The Underdog Narrative and Future Outlook
The conversation concludes with a reflection on Google's journey in AI and its future.
- Underdog Story: The initial phase at Google felt like an "underdog story," despite infrastructure advantages, as the team was building products like AI Studio with limited users and revenue in the early stages of Gemini.
- Catching Up and Leadership: There was a period of "catching up" in LLMs, with an honest assessment of being "nowhere near" state-of-the-art. However, through innovation and a strategic approach, Google has now reached a leadership position.
- Unique Advantages: Google's size and infrastructure are seen as advantages that enable unique capabilities and scale.
- Innovation as a Differentiator: The team has innovated in technology, models, processes, and operational methods, creating a unique approach.
- Building Intelligence the Right Way: The ultimate goal remains to build intelligence "the right way," with all minds and innovation focused on this objective.
- Exciting Future: The next six months are anticipated to be as exciting as the previous ones, with continued progress and innovation. The journey is ongoing, with a focus on achieving the ambitious goal of AGI.
AI summaries can miss context or contain errors. Check important details against the original video.