Living Canvas, a web-based puzzle game powered by Generative AI

Google for DevelopersAbout 4 min readMay 29, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI-powered game design
  • Image analysis and property mapping
  • Text-to-gameplay function transformation
  • Image and video generation models (Gemini, Imagen, Veo)
  • Firebase Studio for prototyping
  • Angular and Phaser JS for front-end development
  • Vertex AI for accessing Google's AI models
  • Multimodal requests
  • Art style transfer
  • Dynamic level generation

Living Canvas: An AI-Powered Puzzle Game

Game Overview

"Living Canvas" is a web-based puzzle game where users draw objects and write commands to interact with the game world. AI analyzes the drawings, assigns gameplay properties, and upgrades the graphics. The game leverages Angular, Phaser JS, Vertex AI, Gemini, Imagen, and Veo.

Core Gameplay Mechanics

  1. Drawing and Image Analysis: Users draw on a canvas. The drawing is sent to the server for analysis.
  2. Property Assignment: AI identifies the drawing and assigns relevant gameplay properties (e.g., drawing a fire assigns the "burning" property).
  3. Object Interaction: The drawing becomes an interactive object in the game world, with its assigned properties affecting gameplay.
  4. Text-Based Commands: Users can write commands (e.g., "douse all the fires") that the AI interprets and executes in the game.
  5. Graphics Upgrading: User drawings are upgraded to higher-fidelity versions using image-generation models.

Technology Stack

  • Front-end: Angular for UI, Phaser JS for game engine.
  • Back-end: Node.js running on Firebase App Hosting.
  • AI Models: Gemini, Imagen, and Veo, accessed via Vertex AI on Google Cloud.
  • Development Environment: Firebase Studio for easy code exploration and integration.

Image Generation Models

  1. Gemini Flash: Used for image analysis and image generation.
  2. Imagen: Used for high-fidelity static image generation. Called via Vertex AI.
  3. Veo: Used for video generation. The generated video is then used to create an animated sprite.

Walkthrough Examples

  1. Melting Ice: The user draws a fire, which is recognized by Gemini. The fire object, with the "burning" property, melts an ice block, allowing the key to reach the lock.
  2. Heart Animation: The user draws a heart. Imagen generates a high-fidelity image of the heart. This image is then sent to Veo to generate an animated video. Frames from the video are used to create an animated sprite in the game.
  3. Spell Casting: The user writes "douse all the fires." Gemini analyzes the text and transforms it into a JSON command that the game engine executes, extinguishing all fires in the game.

AI Implementation Details

  1. Image Analysis with Gemini:
    • The server receives the user's drawing as Base64 data.
    • Gemini first checks if the image matches any predefined object mappings (e.g., fire = burning property).
    • If no match, Gemini provides a text description of the image.
    • This description is sent back to Gemini with a list of possible properties, and Gemini determines which properties are appropriate.
  2. Text Command Processing with Gemini:
    • The user's text command is sent to Gemini.
    • Gemini transforms the command into a JSON format that the game engine can understand and execute.
    • The verbs are predefined functions, and the objects are determined by what exists in the game world.
  3. Image Generation with Gemini and Imagen:
    • The user's drawing and a text description are sent to the image backend.
    • A description of the desired art style (realistic, cartoon, pixelated) is also included.
    • Vertex AI is used to call the Imagen image-generation backend.

Video Generation with Veo

  • Veo is used to generate short video clips from images.
  • The process is time-consuming (up to a minute).
  • FFmpeg is used to extract individual frames from the video.
  • These frames are used to create an animation loop in Phaser JS.

Dynamic Level Generation

  • Collision masks (black and white visuals) are designed to define the level layout.
  • A generation script sends the mask to Gemini, along with a text prompt to generate a game-level background graphic that matches the mask.
  • This process is done for each of the three visual styles (realistic, cartoon, pixelated).

Notable Quotes

  • "Truly great breakthroughs in game design are about creating experiences which were not possible before new technologies existed."

Conclusion

"Living Canvas" demonstrates how AI can unlock new possibilities in game design, enabling dynamic property assignment, natural language interaction, and high-fidelity graphics generation. The game leverages Google's AI models (Gemini, Imagen, Veo) via Vertex AI, along with Firebase Studio for rapid prototyping. The project highlights the potential of AI to create innovative and engaging gaming experiences.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.