Learn to Build with Gemini Nano-Banana (Gemini 2.5 Flash Image)

Google for DevelopersAbout 4 min readSep 6, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Gemini 2.5 Flash Image (Nano Banana), Google AI Studio, Gemini Developer API, Google Gen AI SDK, Image Generation, Image Editing, Multiple Input Images, Photo Restoration, Colorization, Conversational Image Editing, Chat Sessions, Prompt Engineering.

Prototyping in AI Studio

  • Main Point: Google AI Studio is the recommended starting point for developers to experiment with Nano Banana.
  • Process:
    1. Go to ai.studio and sign up with a Google account.
    2. Select Nano Banana in the model picker.
    3. Enter a prompt (e.g., "restore and colorize the image").
    4. Run the model and download the restored image.
    5. Continue chatting for multiple image editing steps.

Project Setup

  • Main Point: To use Nano Banana in your code, you need an API key, billing setup, and the Google Gen AI SDK.
  • Steps:
    1. Create API Key: In AI Studio, click "Get API Key," then "Create API Key." Choose a Google Cloud project (or create one) and copy the generated key.
    2. Set up Billing: Click the billing setup link in AI Studio. Nano Banana costs approximately $0.04 per image. $1.00 can create roughly 25 images. Check the official pricing page for up-to-date info.
    3. Install SDK: Use pip install -u google-genai Pillow python-dotenv to install the necessary Python packages.

Image Creation

  • Main Point: Generating images involves setting up the client, creating a prompt, and calling the generate_content method.
  • Code Example (Python):
    import google.generativeai as genai
    from PIL import Image
    from io import BytesIO
    from dotenv import load_dotenv
    import os
    
    load_dotenv()
    genai.configure(api_key=os.getenv('GEMINI_API_KEY'))
    model = genai.GenerativeModel('gemini-2.5-image-preview')
    
    prompt = "create a photorealistic image of an orange cat with green eyes sitting on a couch"
    responses = model.generate_content(prompt)
    
    for part in responses.parts:
        if hasattr(part, 'text'):
            print(part.text)
        else:
            im = Image.open(BytesIO(part.inline_data.data))
            im.save('cat.png')
    
  • Explanation:
    • The code loads the API key from an environment variable.
    • It initializes the Gemini model with the model ID (gemini-2.5-image-preview).
    • It calls generate_content with the prompt.
    • It iterates through the response parts, saving image data to a file.
  • Model ID: The model ID can be copied from AI Studio to ensure it's up-to-date.

Image Editing

  • Main Point: Editing images requires loading an image and including it in the contents argument of the generate_content call.
  • Code Example (Python):
    prompt = "using the image of the cat. Create a street level view of the cat walking in New York City and then some more details we want."
    image = Image.open('cat.png')
    responses = model.generate_content([prompt, image])
    
    for part in responses.parts:
        if hasattr(part, 'text'):
            print(part.text)
        else:
            im = Image.open(BytesIO(part.inline_data.data))
            im.save('cat2.png')
    
  • Explanation:
    • The code loads an existing image using Image.open.
    • The contents argument is now a list containing the prompt and the image.

Working with Multiple Input Images

  • Main Point: Multiple images can be used as input by including them in the contents list.
  • Code Example (Python):
    prompt = "make the girl wear this T-shirt. Leave the background unchanged."
    image1 = Image.open('girl.png')
    image2 = Image.open('tshirt.png')
    responses = model.generate_content([prompt, image1, image2])
    
    for part in responses.parts:
        if hasattr(part, 'text'):
            print(part.text)
        else:
            im = Image.open(BytesIO(part.inline_data.data))
            im.save('girl_with_tshirt.png')
    
  • Example: The example shows how to make a girl wear a specific T-shirt while keeping the background unchanged.

Photo Restoration and Colorization

  • Main Point: Nano Banana excels at restoring and colorizing old photos with a simple prompt.
  • Prompt Example: "restore and colorize this image. This is a photograph from 1932."
  • Result: The model can produce impressive results in restoring old photos.

Conversational Image Editing (Chat Sessions)

  • Main Point: Chat sessions allow for iterative image editing by maintaining a conversation history with the model.
  • Process:
    1. Create a chat session using client.chats.create with the model ID.
    2. Send messages using chat.send_message, including images and prompts.
    3. Display or save the generated images.
    4. Continue sending messages to refine the image further.
  • Code Example (Python):
    chat = model.start_chat()
    response = chat.send_message(["change the cat to a bengal cat", Image.open('cat.png')])
    # Display or save the image from response
    
    response = chat.send_message("The cat should wear a funny party hat")
    # Display or save the image from response
    

Best Practices for Prompt Engineering

  • Be Specific: Provide detailed context and intent in your prompts.
  • Iterate and Refine: Continuously adjust prompts based on the results.
  • Step-by-Step Instructions: Use step-by-step instructions for complex tasks.
  • Positive Framing: Frame prompts positively (e.g., "an empty, deserted street" instead of "no cars").
  • Control the Camera: Use terms like "wide-angle shot" or "macro shot" to influence the image's perspective.

Conclusion

Nano Banana offers powerful image generation and editing capabilities accessible through the Gemini Developer API. By using Google AI Studio for prototyping, setting up the API correctly, and following best practices for prompt engineering, developers can effectively integrate Nano Banana into their applications for tasks ranging from simple image creation to complex conversational editing and photo restoration. The key is to be specific with prompts, iterate on results, and leverage the model's ability to handle multiple input images and maintain context in chat sessions.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.