Release Notes: Building Gemini's Coding Capabilities

Google for DevelopersAbout 7 min readJun 17, 2025Watch original
THE SUMMARYAI-generated

Gemini Coding Capabilities: A Deep Dive

Key Concepts:

  • Competitive Programming vs. Real-World Coding: The limitations of using competitive programming benchmarks (e.g., LeetCode) as the sole measure of coding model effectiveness.
  • Repo Context: The importance of training models to understand and operate within the context of large code repositories.
  • Vibe Coding: The concept of using natural language to generate functional applications or code snippets, even without deep programming knowledge.
  • Model Fundamentals: Ensuring the underlying model architecture and training processes are sound before attempting specific improvements.
  • Instruction Following: The ability of a model to accurately interpret and execute instructions, crucial for coding tasks.
  • Agentic Coding: Utilizing AI agents to autonomously navigate codebases, identify issues, and implement solutions.
  • Evals (Evaluations): Methods for assessing the performance and capabilities of coding models, including real-world testing and proxy metrics.
  • Model Style/Personality: The impact of a model's tone and presentation on user experience and acceptance.

1. The Journey to Great Coding Models

  • Initial Goals (One Year Prior):
    • Competitive Programming: While useful for initial evaluations (e.g., HumanEval), it doesn't reflect real-world developer tasks.
    • LMSYS Leaderboard: Not representative of day-to-day coding activities.
    • Code Completion: Too limited in scope compared to the potential of coding models.
  • Shifting Focus: The team realized the need to focus on tasks more reflective of real-world developer workflows, such as working within large repositories and making multi-file edits.
  • Fundamentals First: Before making specific changes, the team focused on ensuring the underlying model architecture and training processes were sound. This involved aligning the goals of different teams working on various aspects of Gemini.
  • Competitive Programming Limitations: Competitive programming tasks are self-contained and don't require understanding large codebases or debugging complex issues spread across multiple files.

2. Key Ingredients of a Great Coding Model

  • Data Methodology: Data and methodology are key, shifting to meet each other accordingly.
  • Repo Context: Training models to understand and operate within the context of large code repositories is crucial.
  • Multi-File Edits: Focusing on enabling models to make larger, more complex changes that developers would typically spend hours on.
  • Developer Workflow Inspiration: Drawing inspiration from how developers use models and anticipating future needs.
  • Addressing Developer Pain Points: Identifying and addressing the challenges developers face throughout the entire software development process, not just the code editing part.

3. Vibe Coding and the Future of Development

  • Expanding Access: The goal is to empower people without extensive programming knowledge to perform basic tasks using code.
  • Beyond Web Apps: Vibe coding can extend to various applications, as long as users can validate the final output.
  • Timing and Bets: Getting the timing right and making the right bets on future trends is a constant challenge.
  • Why Not Earlier?: The models were likely capable, but there wasn't enough focus on developer workflows. The belief that users would want to generate entire web apps from natural language wasn't widespread.
  • Andrej Karpathy's Influence: Credited with popularizing the term "vibe coding," helping to define and promote the concept.

4. Gemini's Coding Pillar and Interconnected Capabilities

  • Interconnectedness: Improvements in coding capabilities are a result of a team effort across all of Gemini, including capability-focused and horizontal-focused efforts.
  • Coding for Non-Coding Problems: Using code as an intermediate step to solve problems outside the coding sphere (e.g., converting word problems into code for reasoning).
  • Code Generation for Every Query?: Exploring the possibility of automatically generating code to answer user queries, even if not explicitly requested.
  • Tax Calculation Example: Suggests that even seemingly simple natural language requests (e.g., "give me tips on how to do taxes") could be addressed by generating code to perform the underlying calculations.
  • "Bring Down the Cost of Eggs" Eval: A pseudo-evaluation concept where the model is tasked with solving real-world problems that benefit society, such as lowering the price of eggs by analyzing publicly available data.

5. Evaluating Coding Models: Challenges and Future Directions

  • Avoiding Over-Optimization: The goal is to tackle the core, fundamental challenges in coding, rather than focusing on quick wins.
  • Real-World Testing: A/B testing in real-world scenarios is the most representative evaluation method, but it's often impractical.
  • Proxy Metrics: Pragmatic trade-offs are made to find easier-to-measure proxies for real-world performance.
  • Breadth of Use Cases: The challenge is to build capabilities that generalize across all the different ways people are using code models.

6. Leveraging Google's Internal Expertise

  • 100,000+ Engineers: Access to a vast pool of talented and opinionated engineers provides invaluable feedback on code quality.
  • Nuanced Tastes: Live feedback from professional developers helps address the nuanced tastes and preferences that are often missed in standard evaluations.
  • Jeff Dean and Emma Examples: Using feedback from highly respected engineers like Jeff Dean and Emma to identify areas for improvement and define new tiers of capability.
  • Internal vs. External Feedback: The feedback from internal Googlers and the external developer ecosystem is generally similar in terms of what matters for model capabilities.

7. Addressing AI Skeptics

  • Skeptics as a Guide: Skeptics provide valuable insights into areas where the model needs improvement.
  • Targeted Improvements: The goal is to make the model better at the specific tasks that skeptics genuinely need and care about.
  • Winning Over Skeptics: Focus on understanding what's holding them back and demonstrating the model's capabilities in those areas.
  • Language and Framework Considerations: Optimizing the training mixture to ensure adequate representation of different languages and frameworks based on user needs.

8. The Future of Programming Languages

  • Python and JavaScript Dominance: The possibility that Python and JavaScript could become the dominant programming languages due to the models' proficiency in those languages.
  • New Language Opportunities: The potential for a new programming language to emerge specifically designed for the AI age.
  • Internal Language Eval: Creating an internal programming language and evaluating the model's ability to learn and use it based on its specification.

9. Long Context and Agentic Coding

  • Problem vs. Solution: Organizing approaches based on the problem (working with complex codebases) and the solution strategy (long context vs. agentic coding).
  • Long Context Approach: Loading the entire codebase into the model's context window and solving problems in a single step.
  • Agentic Approach: Using AI agents to autonomously navigate the codebase, search for information, and implement solutions.
  • Mixing and Matching: Combining long context and agentic approaches for optimal results.
  • Non-Human Strategies: The potential for models to discover novel, non-human strategies for solving coding problems.
  • Interpretability: The importance of ensuring that the model's changes are interpretable and understandable, even if it uses unconventional strategies.

10. The Future Direction of Gemini's Coding Capabilities

  • Continuous Improvement: The focus is on setting the right benchmarks and continuously improving the model's capabilities.
  • Tool Calling Functionality: Addressing issues with tool calling functionality, particularly in the context of codebases.
  • Fine-Tuning User Interaction: Improving the smoothness and intuitiveness of user interactions with the model.

11. Model Style and Personality

  • Catering to High Taste: The importance of creating models that generate visually appealing and well-designed outputs.
  • Visual Improvements: Targeted methodologies for improving the visual layouts of model-generated UIs.
  • Tone and Personality: The impact of a model's tone and personality on user experience and acceptance.
  • Balancing Professionalism and Approachability: Tailoring the model's style to different user groups, such as professional developers and beginners.

12. Aha Moments and Initial Experiences

  • Platformer Game Example: Danny's experience with a platformer game where the model successfully identified and fixed unreachable platforms across multiple levels.
  • Surprise and Excitement: The realization that the 2.5 Pro model was exceeding expectations and achieving significant improvements in coding capabilities.
  • Early Generative Models: Danny's early work on generative models of source code during his PhD and postdoc.
  • Structured Models: Building models that incorporate human knowledge and capture basic programming rules.
  • Connie's Journey: Connie's initial belief in the potential of coding models and her personal aha moment when she started using Cursor.

13. Why Not a Code-Specific Model?

  • Domain-Specific Models: While useful for narrow product-specific tasks (e.g., code completion), they lack the world knowledge and reasoning abilities needed for broader coding tasks.
  • Interconnectedness: Code is increasingly interconnected with other aspects of software development and requires access to diverse information sources.
  • Generalist Approach: The focus is on building a generalist model that excels at both coding and other tasks, allowing for seamless integration and improved overall performance.

Synthesis/Conclusion:

The development of Gemini's coding capabilities has been a journey of continuous learning and adaptation. The team has moved beyond traditional benchmarks to focus on real-world developer workflows, leveraging the expertise of Google's internal engineers and addressing the needs of both professional developers and novice coders. The future of coding models lies in their ability to seamlessly integrate with other AI capabilities, understand complex codebases, and even discover novel solutions to coding problems. While challenges remain, the team is confident in their ability to continue pushing the boundaries of what's possible with AI-powered coding.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.