Deploying A GPU-Powered LLM on Cloud Run

By Google Cloud Tech

Share:

Key Concepts AI Model Deployment, Serverless Computing, GPU Acceleration, Google Cloud Run, Ollama, Gemma 3 270M, Dockerfile, Independent Scaling, Fine-tuning, Cold Start Times, Nvidia L4 GPU, Concurrency, Max Instances, LLM (Large Language Model), Instruction-Tuned, Quantization.

Introduction and Rationale for Decoupled LLM Deployment

The video introduces the fascinating concept of an AI model as a "file of numbers" that gains functionality upon internet deployment. The core objective is to demonstrate how to deploy a cutting-edge language model with GPU acceleration, completely serverless, and globally scalable, using just a few commands on Google Cloud. This is the first part of a three-part series focused on deploying an open-source model to serve as the "brain" of an AI agent.

A key argument for deploying a separate, decoupled LLM, rather than relying solely on pre-existing models, is the ability to scale it independently. This separation is crucial for cost management, especially when utilizing GPU resources. Additionally, it provides the flexibility to fine-tune a more specific model for targeted use cases. An illustrative example provided is fine-tuning an LLM to specialize in information about a zoo, enabling it to act as an expert zoo tour guide, answering questions about resident animals.

Building the LLM Service ("The Brain")

The initial focus of this video is on building the "brain" of any AI agent: the model itself. The chosen open-source model is Google's new tiny Gemma model, specifically Gemma 3 270M. This model will be deployed as a scalable, GPU-accelerated service using Google Cloud Run. This dedicated service for running the open model is designed to be separate, allowing for independent scaling, which is particularly beneficial given the higher cost of GPU resources compared to CPU resources. Ollama is selected as the framework to serve Gemma.

Ollama and Dockerfile Breakdown

The "magic" of serving the LLM happens within the Dockerfile, which outlines the steps to create the container image.

  1. Base Image: The process begins by starting from the official ollama/ollama Docker image. Ollama is highlighted as an LLM-serving framework that simplifies the deployment of open models.
  2. Model Selection and Characteristics: The series utilizes Google's Gemma 3 270M model. This model is characterized as:
    • Compact: Possessing only 270 million parameters.
    • Energy-efficient: Designed for lower power consumption.
    • Instruction-tuned: Optimized for following instructions effectively.
    • Production-ready quantization: Built-in optimization for efficient inference. These features make it ideal for handling isolated tasks while maintaining reasonably low costs.
  3. Environment Variables:
    • OLLAMA_HOST=0.0.0.0: Configures Ollama to accept requests from any IP address, not just the local machine.
    • OLLAMA_KEEP_ALIVE=-1: This is a critical optimization that instructs Ollama to never unload the model from the GPU's memory. This significantly reduces latency and makes subsequent requests much faster by eliminating the overhead of reloading the model.
  4. Model Baking: The most crucial step is RUN ollama pull $MODEL (where $MODEL refers to Gemma 3 270M). This command bakes the model's weights directly into the container image. The advantage of this approach is that when a new instance of the service starts up, the model is already present, drastically slashing "cold start times" – the delay experienced when a new instance needs to download and load the model.

Deploying with Google Cloud Run

The deployment of the containerized Ollama service with Gemma is executed using a single gcloud run deploy command, which, despite its length, focuses on several key flags:

  • Service Name: ollama-gemma3-270m-gpu identifies the deployed service.
  • Hardware Request: --gpu 1 --gpu-type=nvidia-l4 requests one Nvidia L4 GPU. The L4 GPU is specifically noted as a "beast for inference," offering both speed and cost-effectiveness for AI inference tasks.
  • Resource Allocation: --memory=16Gi --cpu=8 allocates 16 gigabytes of system memory and 8 CPU cores. This ensures sufficient resources to support the GPU and manage the data flow in and out of the service.
  • Concurrency: --concurrency=4 is a vital performance knob. It configures each instance of the service to handle up to four requests simultaneously, thereby keeping the GPU busy and maximizing throughput.
  • Cost Control: --max-instances=3 sets a cap on the maximum number of instances that can spin up. This is a crucial measure for controlling costs and preventing unexpected bills.

Upon executing the command, Google Cloud Build is triggered. It takes the Dockerfile, builds the container image, and deploys it to Cloud Run. This entire process typically takes about five minutes, primarily due to the time required for downloading the model and provisioning the GPU. Once complete, the command provides a URL, signifying that the service is live and ready to serve requests globally.

Conclusion and Next Steps

The video concludes by summarizing the achievement: a powerful open-source language model has been successfully deployed on a serverless platform with dedicated GPU hardware, ready to scale and serve requests from anywhere in the world. This constitutes the "brain" of the AI agent. However, the model cannot yet interact with users. The next video in the series will focus on building the "face" of the application – an ADK agent – and connecting it to this newly deployed, GPU-powered backend.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video