The Smokejumpers: Scaling Gemini’s serving infrastructure

Google for DevelopersAbout 5 min readJan 27, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini Scaling: The challenges and processes involved in deploying Gemini models (Pro, Flash, and others) across Google’s infrastructure to billions of users.
  • TPU Infrastructure: The role of Tensor Processing Units (TPUs) – Google’s custom AI accelerator – in efficiently serving large language models.
  • Smokejumpers: A specialized, nimble team within Google responsible for rapid deployment and stabilization of LLM serving infrastructure.
  • Full Stack AI: Google’s integrated approach to AI development, encompassing model training, hardware (TPUs), and serving infrastructure.
  • Model Serving: The process of making trained AI models available for use, including capacity allocation, latency optimization, and cost management.
  • MoE (Mixture of Experts): A model architecture utilizing multiple "expert" networks to improve performance and efficiency.
  • Embedding Models: Models that represent data as vectors, enabling semantic search and improved context understanding.
  • Cache Hits: The rate at which requested data is found in a cache, impacting latency and cost.

Scaling Gemini: Infrastructure, Teams, and Challenges

The conversation centers around the immense undertaking of scaling Gemini models – Pro, Flash, and future iterations – to serve billions of users across Google’s product ecosystem. Ema Taropa, a Google Fellow, details the evolution of this process, starting from initial experimentation in late 2022 to the current state of multiple launches and a robust pipeline. The core theme is that while model training is challenging, making those models available and efficiently serving them is equally, if not more, complex.

From Initial Experimentation to Global Deployment

Taropa describes a phased approach beginning in November/October 2022 with a focus on rapid experimentation. This involved figuring out how to route requests between different model types, allocate capacity, and initially roll out access to trusted testers. This evolved into the Gemini program, identifying available capacity across the company, initiating the first Mixture of Experts (MoE) program, and ultimately developing large, fast, and cost-effective models. A common thread throughout this evolution has been leveraging Google’s existing expertise in scaling systems, continuously reducing costs, and improving speed. The goal is to achieve “best in class models and best in class distributed infrastructure.”

The Illusion of the “Easy Button” & Real-World Challenges

The initial expectation of a simple “easy button” for deployment – where a trained model is instantly available across all data centers – is quickly dismissed. While the model itself is a significant achievement, the infrastructure work required to serve it is substantial. Key challenges include:

  • Capacity Allocation: Determining how much computational resources to dedicate to training versus serving.
  • Cost Management: Balancing performance with cost-effectiveness.
  • Traffic Prediction & Adaptation: The unpredictable nature of user traffic, query volume, and application usage patterns, requiring constant monitoring and adjustment. Even with careful planning, unexpected costs and performance bottlenecks arise.
  • Trade-offs: Balancing global optimization with individual user experience (e.g., cache hit rates).

The Role of Smokejumpers: Rapid Response & Ownership

The Smokejumpers team, conceived by Taropa and Ben Treynor, is a critical component of this scaling process. Inspired by wildfire firefighters, the team is designed to be nimble and focused on rapid deployment of LLM serving infrastructure. Key characteristics of the Smokejumpers include:

  • Cross-Functional Composition: Members come from various backgrounds (SRE, engineering, PM) with a shared aptitude for intense problem-solving.
  • Strong Ownership: A sense of responsibility for the end-to-end deployment process.
  • Pressure Cooker Environment: The team thrives in a high-pressure, fast-paced environment.
  • Esprit de Corps: A strong team spirit fostered through collaboration and shared challenges.

TPU Infrastructure & Co-Development

Google’s Tensor Processing Units (TPUs) are fundamental to Gemini’s serving infrastructure. Taropa highlights the advantages of Google’s vertically integrated approach to chip development, led by Amin Vahdat. This allows for:

  • Deep Collaboration: Close collaboration between the TPU development team and the serving infrastructure teams.
  • Rapid Iteration: The ability to quickly adapt TPU roadmaps to meet the specific needs of LLM serving.
  • Institutional Knowledge: A long-standing team with deep expertise in hardware and software optimization.

3.0 Pro vs. 2.5: Lessons Learned & Continued Optimization

The launch of Gemini 3.0 Pro benefited from lessons learned during the 2.5 Pro deployment. While 3.0 Pro is a significantly improved model, 2.5 Pro continues to see substantial usage, highlighting the importance of maintaining older models. Ongoing optimization efforts focus on:

  • Latency Reduction: Improving response times.
  • Cost Optimization: Reducing the computational cost of serving.
  • Cache Hit Rates: Increasing the efficiency of data retrieval.

The Cache Hit Rate Dilemma & Global vs. Local Optimization

A key challenge discussed is optimizing cache hit rates. Developers have reported lower cache hit rates for Gemini compared to other providers. Taropa explains that while Google has robust caching infrastructure for systems like Spanner and Search, adapting it to the specific requirements of LLM serving is ongoing. The challenge lies in balancing global optimization with the unique usage patterns of different applications (e.g., Gmail vs. Antigravity). The goal is to improve cache hit rates to reduce costs and improve performance for all users.

The Human Element & Team Culture

Beyond the technical challenges, the conversation emphasizes the importance of the human element. Taropa shares a personal anecdote about a bike accident and the outpouring of support from the team, illustrating the strong bonds and collaborative spirit within the organization. He emphasizes that the greatest satisfaction comes from the friendships forged during this work. The MK (Google’s main campus) is described as fostering this collaborative environment.

Flash Model & Future Directions

The launch of Gemini Flash is lauded as a success, offering a high-quality model at an efficient cost. Taropa notes that while Flash is powerful, it’s important to balance its capabilities with the needs of a broad user base. The discussion also touches on the potential of embedding models, particularly for applications like Workspace, where they can enhance context understanding and improve the overall user experience.

Conclusion

Scaling Gemini is a complex, iterative process that requires a combination of cutting-edge technology (TPUs), specialized teams (Smokejumpers), and a collaborative, full-stack approach. While challenges remain – particularly around cost optimization and cache hit rates – Google is committed to making these powerful models accessible to billions of users. The success of this endeavor is not only a testament to technical innovation but also to the strong team culture and dedication of the individuals involved.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.