Build an AI app that watches videos using Gemini
By Google Cloud Tech
Key Concepts
- Gemini 2.5: A multi-modal AI model capable of processing and generating text, audio, and video content from a single API call.
- Multi-modal AI: Artificial intelligence that can understand and process information from multiple modalities (e.g., text, image, audio, video) simultaneously.
- Prompt Engineering: The process of designing and refining prompts to guide an AI model to produce desired outputs.
- Imagen: A Google AI model specifically for generating images from text prompts.
- Veo: A Google AI model mentioned for its capabilities in video generation/processing.
- Tokens: The basic units of text or data that AI models process, used for pricing and capacity measurement.
- API Call: A request made to an application programming interface to perform a specific function.
Revolutionizing Video Content Processing with Gemini 2.5
The video introduces a paradigm shift in how applications can "watch" and process video content, moving from a complex, multi-step pipeline to a simplified, single API call using Google's Gemini 2.5.
The Traditional vs. Gemini 2.5 Approach
Traditionally, processing video content for tasks like summarization would involve a cumbersome pipeline:
- Audio Extraction: Pulling out the audio track from the video.
- Speech-to-Text Service: Converting the audio into a text transcript.
- OCR Model: Optical Character Recognition for any slides or on-screen text.
- Summarizer: Processing the combined text to generate a summary.
This entire, "serious pipeline" is now condensed into a single API call to Gemini 2.5, a multi-modal AI model. Ayo, a Developer Relations Engineer at Google, demonstrates this capability.
Demo Application: YouTube URL to Blog Post
A practical application showcases Gemini 2.5's power:
- Input: A YouTube URL.
- Process: The application makes an API call to Gemini.
- Output: A complete blog post and a header image derived from the video's content.
- Example: The demo used one of Martin's (the host's) old videos, successfully generating relevant content.
Technical Breakdown of the Application
The application's core logic is surprisingly simple, relying on just two API calls:
-
Blog Post Generation:
- Files:
index.htmldefines the input field (youtube_link), andapp.pycontains the main logic. - Function:
generate_blog_post_texttakes the YouTube link and a chosen model name. - Process: It constructs a request containing the YouTube link and a text prompt, sends it to Google's GenAI (Generative AI), and receives the blog post text.
- Key Detail: Gemini 2.5 directly "watches" the video and "listens" to the audio within this single call; no prior transcript or audio extraction is needed.
- Files:
-
Header Image Generation:
- Function:
generate_imageis responsible for creating the visual component. - Input: The title of the generated blog post (e.g., "Taming the Serverless File System; A Practical Guide").
- Process: It obtains a specific prompt for image creation, sends this prompt along with settings to the Imagen model, and receives a single PNG image.
- Display: The image is then encoded to base64 for display on the web page.
- Function:
The entire application, therefore, consists of these two API calls plus minimal supporting code.
The Power of Prompt Engineering
The "magic" behind Gemini's output largely resides in the prompt.
- Prompt Structure: The prompt, returned by the
get_blog_gen_promptfunction, includes:- A defined persona for the AI to adopt.
- Top-level instructions.
- Detailed instructions formatted in markdown.
- Prompt Engineering Process:
- Ayo explains that achieving the perfect prompt is an iterative process.
- He started with an AI-generated prompt, then refined it by "adding background context, defining its role, and adding clear layout rules until it finally worked." This highlights that initial prompts often require significant iteration.
Cost Considerations
Google provides clear metrics for pricing:
- Token Count: One second of video is approximately 300 tokens. A full minute of video is roughly 18,000 tokens.
- Pricing Model: Using Gemini 2.5 Flash, priced at $0.30 per million tokens.
- Estimated Cost: Processing one minute of video would cost about half a US cent, after utilizing any daily free quota.
- Cost Optimization: Switching to low-resolution video processing can reduce costs by two-thirds.
Video Source Flexibility
The API is not limited to YouTube URLs:
- Users can upload MP4 files directly.
- The API can be pointed to a video stored in Cloud Storage.
- Regardless of the source, Gemini 2.5 processes the audio and video data in that single, efficient call.
Versatility and Broader Applications
The demonstrated app is not just a demo but a "pattern" for leveraging multi-modal AI:
- Prompt Customization: By simply changing the prompt, the same underlying mechanism can generate different outputs, such as bullet-point summaries or a 10-question quiz from the video content.
- Dynamic Prompts: Prompts can be stored in external text files or databases, allowing developers to modify application behavior without altering code or deploying new versions.
- Multi-modal Output: Gemini 2.5, along with Veo for video, can also output text, audio, and video formats. This opens possibilities like:
- Converting a blog post directly into an audio script.
- Transforming meeting recordings into video highlight reels.
Future Enhancements
The discussion briefly touches upon the potential for user and AI collaboration in editing generated content, such as a blog post. While feasible, this would require additional code and is suggested for a future development.
Conclusion
Ayo emphasizes that "multi-modal models like Gemini can do a lot with just one API call, and there's a lot of power in the prompt." The ability of Gemini 2.5 to process video and audio content directly from a URL or file, combined with the flexibility of prompt engineering, significantly simplifies complex content generation pipelines, offering powerful and cost-effective solutions for developers.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development