Gemini 2.0 blew me away - The future of Multimodal Model

AI JasonAbout 6 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini 2.0 Experimental Model: Google's multimodal model supporting both image understanding and generation.
  • Image Generation & Editing: Using Gemini 2.0 to create and modify images based on text prompts.
  • Multimodal Model: An AI model that can process and generate different types of data, such as text and images.
  • Replicate: A platform for hosting and deploying AI models, including image and video generation models.
  • 1 2.1: An image-to-video model hosted on Replicate, used to generate videos from images.
  • Streamlit: A Python framework for building interactive web applications.
  • API (Application Programming Interface): A set of rules and specifications that software programs can follow to communicate with each other.

Gemini 2.0 Experimental Model: Overview and Capabilities

  • Introduction: Google released Gemini 2.0, a multimodal model capable of both understanding and generating images.
  • Image Understanding and Generation: Users can upload an image and provide a text prompt, and the model will respond with both text and a generated image.
  • Examples:
    • Combining images: Uploading an image of a model and an image of clothes to create a new image.
    • Image extraction: Extracting a passport photo from an image with high fidelity.
    • Animation generation: Generating multiple images in a row to create frames for an animation or GIF.
  • API Availability and Cost: The experimental model is available via API and is significantly cheaper (96% cheaper) than OpenAI's GPT-4o and Cloud.
  • Potential Applications: AI-native Photoshop, GIF makers, and other applications where users can interact with the AI to update images or generate animations.

Testing Gemini 2.0's Image Generation Capabilities

  • Example Use Cases:
    • Image editing: Adding chocolate toppings to an image of ice cream.
    • Visual story generation: Generating a story with an image for each scene.
  • Personal Tests:
    • Changing a flag: Uploading an image of a man with a flag and changing the flag to the USA flag. The result was "extremely good," with minor facial changes.
    • Sketch to 3D render: Converting a sketch into a 3D render with a colorful style.
    • GIF generation: Generating five frames of a 2D pixel game of a dragon monster. The results were "awesome," with some inconsistencies in the last two images.
  • Observations:
    • High-quality image generation on the first shot.
    • Quality decreases with more turns in the conversation (prompting).
    • Good at maintaining character consistency across different images.
  • Conclusion: Gemini 2.0 is impressive and enables more people to do image editing jobs, potentially leading to new types of Photoshop or Canva experiences.

Building a Prototype with Gemini 2.0 API

  • Project Setup:
    • Creating a new project in Cursor.
    • Adding an .env file to store the Gemini API key (obtained from Google AI Studio).
    • Creating a gemini_experimental.py file.
  • Code Implementation:
    • Importing necessary libraries.
    • Creating a Gemini client.
    • Creating a user prompt message (using types.Part).
    • Passing the conversation history to the Gemini 2.0 model.
    • Configuring the response modality to be both text and image.
    • Handling different response types (text or image).
    • Saving the generated image to a file.
  • Testing Image Generation:
    • Running the script to generate an image of a cat.
  • Passing an Image to Gemini:
    • Adding a types.Part from bytes with the image data.
    • Updating the prompt to modify the cat's hair color to red.
    • Running the script again to see the updated image.

Turning Images into Videos with 1 2.1

  • Introduction to Replicate: Replicate is a model marketplace with various AI models for image generation, video generation, and large language models.
  • Using the 1 2.1 Model:
    • Creating a Replicate account.
    • Selecting the "1 2.1 480p" model.
    • Copying the Python code example from the API page.
  • Modifying the Code:
    • Changing the code to read an image from the local disk instead of a URL.
    • Creating a function to open a local image, pass it to the Replicate model, and save the video.
  • Testing the Video Generation:
    • Running the function with the generated cat photo and a prompt ("cat is looking around").
    • Ensuring Replicate is installed (pip install replicate).
    • Verifying the generation of a 5-second video.

Building a Web Application with Streamlit

  • Creating a General Function for Gemini Response:
    • Detecting if it's the first message from the user.
    • Attaching the uploaded image as part of the content.
    • Appending messages and generating a response.
  • Creating a Function for Video Generation:
    • Calling the Replicate model to generate a video.
    • Returning the video path when ready.
  • Creating Helper Functions:
    • Resetting the video state.
  • Utility Functions (in utility.py):
    • Saving binary files.
    • Processing uploaded images.
    • Checking if images are duplicated.
  • Building the GUI with Streamlit:
    • Importing necessary packages and libraries.
    • Setting a title for the app.
    • Defining a list of states to track messages, uploaded images, Gemini-returned images, and 1 2.1-returned videos.
    • Creating a sidebar for users to upload images.
    • Displaying the list of images and updating the state.
    • Creating two tabs:
      • Chat Experience: Allowing users to chat with Gemini to iterate on the image. Displaying chat history and handling the "Send" message logic.
      • Video Generation: Displaying the images returned by Gemini for users to select. Allowing users to select an image and generate a video. Displaying the generated video on the screen.
  • Running the Application:
    • Using streamlit run app.py to start the web app.
  • Testing the Web App:
    • Uploading an image of a bracelet.
    • Prompting Gemini to generate a product shot of a hand wearing the bracelet.
    • Prompting Gemini to change the hand to a black man's hand.
    • Generating a video from one of the generated images with the prompt "a p shot at showcasing the bracelet."

Conclusion and Call to Action

  • Summary: The video demonstrates how to use the Gemini 2.0 API to build a prototype application that combines image generation, editing, and video creation.
  • AI Builder Club Community: Invitation to join the AI Builder Club Community for in-depth API usage, step-by-step rebuilding of the example, tips and tricks for building AI applications, and weekly live coding sessions.
  • Replicate Credits: Offer of $100 free Replicate credits for the first 1,000 AI Builder Club members.
  • Community Benefits: Access to a community of top AI builders for asking questions and sharing learnings.
  • Link in Description: Link to join the AI Builder Club Community provided in the video description.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.