Key Concepts
- Thinking in Gemini: A process where the model can emit additional text in a thinking stage before providing a final answer, allowing for more test-time compute.
- Test-Time Compute: The computational resources a model uses to process a specific request or question.
- Bottlenecks to Intelligence: Limitations in model architecture or training that restrict performance.
- Reinforcement Learning (RL): A training method where the model learns through rewards and penalties.
- Deep Think: A high-budget thinking mode in Gemini 2.5 Pro for complex problems, leveraging deeper and parallel chains of thought.
- Inference Tokens: The units of text processed during the thinking stage.
1. Research Motivation: Unblocking Bottlenecks to Intelligence
- The speaker, Jack, a researcher and tech lead of thinking within Gemini at Google, emphasizes that progress in AI is driven by identifying and solving key bottlenecks.
- Historical Examples:
- Claude Shannon (1948): Shannon's language model was limited by the small amount of data and elementary statistics available. The solution required the digitalization of human knowledge and modern computing.
- Google in the 2000s: En-gram language models were limited by short context due to exponential storage costs. The solution was recurrent neural networks (RNNs) that could store compressed representations of the past.
- RNN Bottleneck: RNNs had a fixed-size state, leading to lossy context representation. The solution was attention mechanisms and transformers, which allowed models to keep all past embeddings and aggregate them on the fly.
- Current Bottleneck (2024): Large language models (LLMs) are trained to respond immediately, limiting the test-time compute they can apply to a specific problem.
- Test-Time Compute Explained: The model translates the request into tokens and processes them through a language model with parallel computation per layer and iterative computation across layers. This computation is fixed.
- Thinking as a Solution: Thinking allows the model to iteratively loop and perform additional test-time compute during a thinking stage, potentially thousands of iterations, before committing to an answer. This process is dynamic, allowing the model to learn how many iterations to apply.
2. Thinking in Gemini: Mechanics and Training
- Mechanics: A thinking stage is inserted into the model's process, allowing it to emit additional text before the final answer.
- Training: The model is trained using reinforcement learning (RL). It receives positive and negative rewards based on whether it solves the task correctly.
- Emergent Behavior: RL training leads to emergent behaviors like hypothesis generation, testing, self-correction, problem decomposition, exploration of multiple solutions, code drafting, intermediate calculations, and tool use.
- Example: The model poses a hypothesis for an integer prediction problem, tests it, rejects it, and tries an alternative approach.
3. Developer Benefits: Capability, Steering, and Efficiency
- Increased Capability: Thinking leads to more capable models by scaling test-time compute. This stacks on top of existing paradigms like pre-training (scaling data and model size) and post-training (scaling human feedback).
- Faster Model Improvement: Investing in pre-training, post-training, and thinking results in a multiplicative effect and faster model improvement.
- Improved Reasoning Performance: There's a correlation between increased test-time compute and improved reasoning performance in math, code, and science.
- Steering Quality Over Cost: Thinking allows for a continuous budget, providing granular control over capability and cost. Thinking budgets are available in Flash and Pro models in the 2.5 series.
4. What's Next: Efficiency and Deeper Thinking
- Efficiency: Focus on making the thinking process more efficient and adaptive, reducing the need for manual tuning. The goal is to prevent overthinking and make models more cost-effective.
- Deeper Thinking: Scaling inference compute further to drive even higher capability.
- Deep Think: A high-budget mode built on top of Gemini 2.5 Pro for very hard problems, leveraging deeper and parallel chains of thought. It enhances performance on multimodal code math problems like the USA Math Olympiad.
- Example: On the USA Math Olympiad, performance increases from negligible to the 50th percentile with 2.5 Pro and to the 65th percentile with Deep Think.
- Application: Open-ended coding tasks that previously took months can now be completed in minutes.
- Example: Gemini vibe-coded the training setup and algorithm from DeepMind's original DQN paper, including an Atari emulator, allowing it to play some games.
5. The Future: Data Efficiency and Deep Contemplation
- The speaker envisions models that can contemplate deeply from a small set of knowledge, similar to Srinivasa Ramanujan, who made significant mathematical contributions with limited resources.
- The goal is to create models that are incredibly data-efficient and can process millions of inference tokens, building up knowledge and artifacts to push the frontier of human understanding.
Synthesis/Conclusion
The presentation highlights the importance of identifying and addressing bottlenecks in AI development. Thinking in Gemini is presented as a solution to the test-time compute bottleneck, enabling models to reason more effectively and tackle complex problems. The future direction focuses on improving the efficiency of the thinking process and scaling it further with "Deep Think," ultimately aiming for models that can deeply contemplate and generate new knowledge from limited information, mirroring the capabilities of the human mind.
AI summaries can miss context or contain errors. Check important details against the original video.





