Build an AI app that watches videos using Gemini

By Google Cloud Tech

Share:

Key Concepts

  • Gemini 2.5: A multi-modal AI model capable of processing and generating text, audio, and video content from a single API call.
  • Multi-modal AI: Artificial intelligence that can understand and process information from multiple modalities (e.g., text, image, audio, video) simultaneously.
  • Prompt Engineering: The process of designing and refining prompts to guide an AI model to produce desired outputs.
  • Imagen: A Google AI model specifically for generating images from text prompts.
  • Veo: A Google AI model mentioned for its capabilities in video generation/processing.
  • Tokens: The basic units of text or data that AI models process, used for pricing and capacity measurement.
  • API Call: A request made to an application programming interface to perform a specific function.

Revolutionizing Video Content Processing with Gemini 2.5

The video introduces a paradigm shift in how applications can "watch" and process video content, moving from a complex, multi-step pipeline to a simplified, single API call using Google's Gemini 2.5.

The Traditional vs. Gemini 2.5 Approach

Traditionally, processing video content for tasks like summarization would involve a cumbersome pipeline:

  1. Audio Extraction: Pulling out the audio track from the video.
  2. Speech-to-Text Service: Converting the audio into a text transcript.
  3. OCR Model: Optical Character Recognition for any slides or on-screen text.
  4. Summarizer: Processing the combined text to generate a summary.

This entire, "serious pipeline" is now condensed into a single API call to Gemini 2.5, a multi-modal AI model. Ayo, a Developer Relations Engineer at Google, demonstrates this capability.

Demo Application: YouTube URL to Blog Post

A practical application showcases Gemini 2.5's power:

  • Input: A YouTube URL.
  • Process: The application makes an API call to Gemini.
  • Output: A complete blog post and a header image derived from the video's content.
  • Example: The demo used one of Martin's (the host's) old videos, successfully generating relevant content.

Technical Breakdown of the Application

The application's core logic is surprisingly simple, relying on just two API calls:

  1. Blog Post Generation:

    • Files: index.html defines the input field (youtube_link), and app.py contains the main logic.
    • Function: generate_blog_post_text takes the YouTube link and a chosen model name.
    • Process: It constructs a request containing the YouTube link and a text prompt, sends it to Google's GenAI (Generative AI), and receives the blog post text.
    • Key Detail: Gemini 2.5 directly "watches" the video and "listens" to the audio within this single call; no prior transcript or audio extraction is needed.
  2. Header Image Generation:

    • Function: generate_image is responsible for creating the visual component.
    • Input: The title of the generated blog post (e.g., "Taming the Serverless File System; A Practical Guide").
    • Process: It obtains a specific prompt for image creation, sends this prompt along with settings to the Imagen model, and receives a single PNG image.
    • Display: The image is then encoded to base64 for display on the web page.

The entire application, therefore, consists of these two API calls plus minimal supporting code.

The Power of Prompt Engineering

The "magic" behind Gemini's output largely resides in the prompt.

  • Prompt Structure: The prompt, returned by the get_blog_gen_prompt function, includes:
    • A defined persona for the AI to adopt.
    • Top-level instructions.
    • Detailed instructions formatted in markdown.
  • Prompt Engineering Process:
    • Ayo explains that achieving the perfect prompt is an iterative process.
    • He started with an AI-generated prompt, then refined it by "adding background context, defining its role, and adding clear layout rules until it finally worked." This highlights that initial prompts often require significant iteration.

Cost Considerations

Google provides clear metrics for pricing:

  • Token Count: One second of video is approximately 300 tokens. A full minute of video is roughly 18,000 tokens.
  • Pricing Model: Using Gemini 2.5 Flash, priced at $0.30 per million tokens.
  • Estimated Cost: Processing one minute of video would cost about half a US cent, after utilizing any daily free quota.
  • Cost Optimization: Switching to low-resolution video processing can reduce costs by two-thirds.

Video Source Flexibility

The API is not limited to YouTube URLs:

  • Users can upload MP4 files directly.
  • The API can be pointed to a video stored in Cloud Storage.
  • Regardless of the source, Gemini 2.5 processes the audio and video data in that single, efficient call.

Versatility and Broader Applications

The demonstrated app is not just a demo but a "pattern" for leveraging multi-modal AI:

  • Prompt Customization: By simply changing the prompt, the same underlying mechanism can generate different outputs, such as bullet-point summaries or a 10-question quiz from the video content.
  • Dynamic Prompts: Prompts can be stored in external text files or databases, allowing developers to modify application behavior without altering code or deploying new versions.
  • Multi-modal Output: Gemini 2.5, along with Veo for video, can also output text, audio, and video formats. This opens possibilities like:
    • Converting a blog post directly into an audio script.
    • Transforming meeting recordings into video highlight reels.

Future Enhancements

The discussion briefly touches upon the potential for user and AI collaboration in editing generated content, such as a blog post. While feasible, this would require additional code and is suggested for a future development.

Conclusion

Ayo emphasizes that "multi-modal models like Gemini can do a lot with just one API call, and there's a lot of power in the prompt." The ability of Gemini 2.5 to process video and audio content directly from a URL or file, combined with the flexibility of prompt engineering, significantly simplifies complex content generation pipelines, offering powerful and cost-effective solutions for developers.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video