AI Engineering with the Google Gemini 2.5 Model Family - Philipp Schmid, Google DeepMind

AI EngineerAbout 7 min readJul 12, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini 2.5 Pro/Flash: Google's multimodal AI models.
  • Multimodality: Ability to understand and generate text, images, audio, videos, and documents.
  • AI Studio: Google's developer platform for testing and experimenting with AI models.
  • API Key: Authentication token required to access Gemini models via API.
  • Google AI SDK: Software Development Kit for interacting with Gemini models.
  • Text Generation: Using Gemini to create text-based content.
  • Token Counting: Estimating the cost of using Gemini based on input and output tokens.
  • Streaming Response: Receiving text generation output in chunks for improved user experience.
  • Chats API: Simplifies managing conversational state with Gemini.
  • Generation Config: Parameters for controlling text generation, including temperature, max output tokens, top P, and top K.
  • File API: Uploading and referencing files (e.g., PDFs, images) for use in prompts.
  • Structured Output: Generating data in a predefined format using Pydantic schemas.
  • Function Calling: Enabling Gemini to request specific actions by returning a function name and arguments.
  • Native Tools: Built-in functionalities like Google Search, code execution, and URL context.
  • Model Context Protocol (MCP): A standard for agent communication and tool integration.
  • Thinking Budget: Controls the number of tokens the model uses for reasoning.
  • Crowning Metadata: Information about the sources used to generate a response.

1. Introduction to Gemini 2.5 and AI Studio

  • The workshop focuses on AI engineering using the Google Gemini 2.0 family, specifically Gemini 2.5 Pro and Flash models.
  • Gemini 2.5 Flash is used due to its free tier availability via API access.
  • Both models are multimodal, capable of understanding and generating text, images, audio, videos, and documents.
  • AI Studio is introduced as a developer platform for testing and experimenting with Gemini models.
  • Actionable Insight: Participants are encouraged to use their personal Gmail accounts for free access and to actively code during the workshop.

2. Setting Up AI Studio and API Key

  • Instructions are provided for obtaining an API key from AI Studio.
  • The process involves creating or selecting a Google Cloud project and generating an API key.
  • The API key needs to be stored as a secret in Google Colab or as an environment variable for local development.
  • A GitHub repository containing workshop notebooks is introduced, with Colab links for easy access.
  • Step-by-Step Process:
    1. Go to AI Studio (AI.dev or AI.studio).
    2. Get API key (top right).
    3. Create or select a Google Cloud project.
    4. Create API key.
    5. Store API key in Colab secrets (Gemini API key) or as an environment variable.
    6. Open the first notebook and run the code snippet.

3. Text Generation and Token Counting

  • The first section of the workshop covers text generation using the Gemini API.
  • It includes examples of generating names for a coffee shop and explaining terms.
  • The generate_content method is used to send prompts to the model and receive responses.
  • The count_token API is introduced to estimate the cost of using Gemini.
  • The workshop explains how to access input tokens, thought tokens, and candidate tokens from the response metadata.
  • Technical Terms:
    • Input Tokens: Tokens in the prompt.
    • Thought Tokens: Tokens used by the model for reasoning.
    • Candidate Tokens: Tokens in the generated response.
  • Data:
    • Gemini 2.5 Flash pricing: $0.10 per 1 million input tokens, $0.40 per 1 million output tokens.
  • Actionable Insight: Understanding token usage is crucial for cost management.

4. Streaming Responses and Chat API

  • The workshop demonstrates how to use streaming responses for a better user experience.
  • The generate_content_stream method is used to receive output in chunks.
  • The Chats API is introduced as a way to simplify managing conversational state.
  • The chat_session object stores the history of the conversation, making it easier to send subsequent messages.
  • The get_history method allows retrieving the complete conversation history for storage or analysis.

5. Generation Configuration and System Instructions

  • The workshop explains how to use the generation_config parameter to control text generation.
  • This includes setting the temperature, max output tokens, top P, and top K.
  • System instructions can be used to guide the model's behavior and ensure compliance with policies.
  • The thinking budget can be set to control the number of tokens the model uses for reasoning.
  • Actionable Insight: Experimenting with different generation configurations can significantly impact the quality and creativity of the output.

6. File API and Document Processing

  • The File API is introduced as a way to upload and reference files for use in prompts.
  • This is particularly useful for working with large documents like books or PDFs.
  • The workshop demonstrates how to upload a book and ask the model to summarize it.
  • The File API can also be used to process PDFs, including extracting information from invoices.
  • Technical Detail: The Gemini API automatically performs OCR on PDFs.
  • Actionable Insight: The File API simplifies working with large documents and reduces the need for manual processing.

7. Multimodality: Image, Audio, and Video Understanding

  • The workshop covers the multimodal capabilities of Gemini, including image, audio, and video understanding.
  • It demonstrates how to upload an image and ask the model to describe it.
  • The workshop also explains how to process audio and video files.
  • Example: Uploading an invoice PDF and asking the model to extract the total amount.
  • Technical Detail: The Gemini API performs OCR on PDFs and provides the image and OCR text to the model.

8. Structured Output and Function Calling

  • The workshop introduces structured output as a way to generate data in a predefined format.
  • Pydantic schemas are used to define the structure of the output.
  • The response_type and response_schema parameters are used to enforce the structure.
  • Function calling is introduced as a way to enable Gemini to request specific actions.
  • The model returns a function name and arguments, which can then be used to call an external API.
  • Example: Defining a weather function and asking the model to call it to get the weather in a specific location.
  • Actionable Insight: Structured output and function calling enable building more complex and integrated AI applications.

9. Native Tools: Google Search, Code Execution, and URL Context

  • The workshop covers the native tools available in Gemini, including Google Search, code execution, and URL context.
  • Google Search allows Gemini to search the web for information and use it in its responses.
  • Code execution allows Gemini to run Python code and return the results.
  • URL context allows Gemini to extract information from a website and use it in its responses.
  • Example: Asking Gemini to find the latest developments in renewable energies using Google Search.
  • Technical Detail: The Google Search tool provides crowning metadata, which indicates the sources used to generate the response.
  • Actionable Insight: Native tools extend the capabilities of Gemini and enable it to perform more complex tasks.

10. Model Context Protocol (MCP) Integration

  • The workshop introduces the Model Context Protocol (MCP) as a standard for agent communication and tool integration.
  • The Google AI SDK now supports native integration of MCP servers.
  • This simplifies the process of connecting Gemini to external tools and services.
  • Example: Connecting Gemini to an MCP weather service to get the weather in a specific location.
  • Actionable Insight: MCP simplifies agent development and promotes interoperability between different AI systems.

11. Conclusion

  • The workshop provides a comprehensive overview of AI engineering with the Google Gemini 2.0 family.
  • It covers a wide range of topics, including text generation, multimodality, structured output, function calling, native tools, and MCP integration.
  • Participants are encouraged to experiment with the Gemini API and explore its capabilities.
  • Feedback and suggestions are welcome to improve the Gemini platform and SDK.
  • Main Takeaway: Gemini offers powerful tools and capabilities for building a wide range of AI applications, and the workshop provides a solid foundation for getting started.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.