Key Concepts
Gemini, Gemini diffusion, Deep Think, generative media models (Imagen, VO, LIA), native audio support, diffusion-based audio generation, model training, post-training, tool use, reasoning scaling, long context, AI evaluation, TPU (Pufferfish/V4), model architecture (Transformers), pre-training, fine-tuning, RL (Reinforcement Learning).
Google IO Announcements and Sergey Brin's Reaction
Sergey Brin expresses excitement about the Google IO announcements, admitting he was unaware of about 30% of them, including the virtual fit feature in Google Search. He emphasizes the importance of shipping the announced features smoothly and ensuring user access. He highlights the phenomenal reception and the need for people to explore and understand the new capabilities.
Focus Areas: Gemini vs. Generative Media
Brin states his primary focus is on the core Gemini text model, believing it will drive self-improvement and advancements in AI science and coding. While acknowledging the "superhuman" nature of generative media models like Imagen and VO, he dedicates most of his time to Gemini. He notes the impressive capabilities of VO, particularly the addition of audio, which elevates video generation from a "gimmick" to a practical tool.
The Impact of Audio in Generative Models
Brin emphasizes the significant impact of adding audio to visual experiences, drawing from his experience with Google Glass. He notes that audio adds a crucial layer of richness and immersion. He observed the VO model training and recognized the transformative effect of audio integration.
Native Audio Support in Gemini and VO
The conversation reveals that native audio support has been present in the base Gemini model for at least a year, but its release was delayed due to prioritization and implementation challenges. VO, on the other hand, uses diffusion for audio generation, similar to its video generation process.
Diffusion as a Powerful Technique
Brin acknowledges the power of diffusion as a technique, referencing Google's early text diffusion experiments. He appreciates the ability to explore different base techniques for various modalities. The results of Gemini diffusion are promising, and he hopes the demo capabilities translate well to real-world performance.
Watching the Model Training Run
Brin describes the process of monitoring model training runs by testing intermediate checkpoints (e.g., 10%, 20% completion). This allows researchers to assess the model's trajectory and identify potential issues early on. He mentions observing this process for both text models and the VO video model.
The Evolution of AI: From Science Fiction to Reality
Brin reflects on past conversations about the future of AI, noting that the current reality is surprisingly close to what was envisioned 10-15 years ago. He acknowledges that while intellectual reasoning about the singularity was possible, witnessing its actual emergence is a different experience. He highlights the unexpected prominence of language models in AI development.
Interpretability and Safety
Brin finds the interpretability of thinking models, which allows for understanding their reasoning processes, to be a positive aspect from a safety perspective. While acknowledging concerns about models "lying," he believes it's a relatively minor issue.
Model Training and System Integration
Brin discusses the architectural similarities between different models, including VO and text language models, with Transformers playing a central role. He notes the increasing importance of post-training, which includes tool use and RL, in shaping model capabilities. Post-training is becoming a more significant part of the overall training process.
Reasoning Scaling with Deep Think
Brin describes Deep Think as a convergence of multiple approaches to reasoning scaling. He emphasizes the potential value of models that can spend extended periods (hours, days, or even months) reasoning about complex problems. He draws a parallel to long context for input, highlighting the non-trivial generalization required for models to handle extended reasoning processes.
AI Evaluation as a Broader Challenge
Brin uses the example of interviewing people to illustrate the broader challenge of evaluation, noting that the AI field is grappling with similar issues. He suggests that the difficulties in AI evaluation reflect the inherent complexities of evaluation in general.
Google as a Startup and the Pace of Innovation
Brin acknowledges the rapid pace of innovation at Google, particularly the acceleration from 2024 to 2025. He emphasizes that Google has always been an AI company, given its focus on large-scale data and its contributions to modern machine learning. He highlights the significance of Gemini 2.5 Pro as a major leap forward and the subsequent launch of Gemini 2.5 Flash, which is positioned as a fast and powerful model.
The Importance of the Science Engine
Brin attributes Google's progress to its "science engine," which drives continuous innovation and model improvements. He emphasizes the importance of the scientific work conducted over the past year in enabling the development of Gemini 2.5 Pro.
Appreciation for Logan and the Team
Brin expresses his appreciation for Logan's hard work in managing customer and partner relationships, addressing technical challenges, and communicating customer needs to the team.
Unboxing the TPU V4 (Pufferfish)
Logan presents Brin with a TPU V4, internally known as "Pufferfish," as a gift. Brin acknowledges the importance of TPUs in their work and expresses hope for acquiring more for the team. He mentions that the presented TPU might be an early sample with potential faults.
Synthesis/Conclusion
The conversation highlights Google's significant advancements in AI, particularly with the Gemini model and generative media. Brin emphasizes the importance of continuous innovation, reasoning scaling, and addressing evaluation challenges. He acknowledges the rapid pace of development and the crucial role of Google's "science engine" in driving progress. The discussion also touches on the architectural similarities between different models, the increasing importance of post-training, and the transformative impact of audio in generative models.
AI summaries can miss context or contain errors. Check important details against the original video.





