THE SUMMARYAI-generated
Key Concepts
- Deepseek AI model
- Cloud Run GPUs
- Ollama (command-line tool)
- Containerization
- Vertex AI API
- Scaling (automatic)
- Google Cloud Storage (GCS)
- Public Preview (Cloud Run GPUs)
Hosting Deepseek AI Model with Cloud Run GPUs
Step 1: Install the Ollama Command-Line Tool
- The first step involves installing the Ollama command-line tool within the Google Cloud Shell.
- Ollama simplifies the process of downloading and running large language models (LLMs).
- It's not limited to Deepseek; it can handle various other models as well.
Step 2: Deploy the Ollama Container as a Cloud Run Service
- The next step is deploying an Ollama container as a new Cloud Run service within the Google Cloud project.
- This container is designed to load LLMs from the internet on demand.
- The command used includes the
--gpuflag, specifically--gpu=1, which instructs Cloud Run to allocate one GPU for each instance of the service. - GPUs are essential for running computationally intensive LLMs like Deepseek.
Step 3: Trigger the Service and Load the Deepseek Model
- After deployment, Cloud Run provides a URL that triggers the newly created service.
- The Ollama command-line tool is then used to interact with this service.
- The
HOSTenvironment variable is set to point to the Cloud Run service's URL. - The
ollama run deepseek-ai/deepseek-llm-7b-chatcommand instructs the Cloud Run service to download the Deepseek model (approximately 5GB) from the internet and load it. - Once loaded, the Deepseek model resides in the Cloud Run service's memory and utilizes the allocated GPU.
Testing the Model
- The Ollama command-line tool initiates an interactive session, allowing users to interact with the Deepseek model.
- A test query, "Why is the sky blue?", is used to verify that the model is functioning correctly.
- The model's response confirms that the Deepseek model is successfully running within the Cloud Run environment.
Alternatives to Hosting Your Own Model
- Vertex AI API: Instead of hosting a model, users can leverage Google's Vertex AI API. This eliminates the need to manage GPUs, Ollama, or other infrastructure components. The Vertex AI client library is used to make API calls.
- Benefits of Hosting: Hosting your own model (like Deepseek, JAMAMA, or a custom-built model) on Cloud Run provides greater control and flexibility.
Architecture and Scaling
- Separation of Concerns: The video highlights an architecture where the application (e.g., an expense reporting application) and the LLM are deployed as separate Cloud Run services.
- Independent Scaling: This separation allows each service to scale independently based on its specific traffic demands.
- Automatic Scaling: Cloud Run automatically spins up more instances of the LLM service during traffic spikes (e.g., Monday mornings for an expense reporting application).
- Scale-to-Zero: During periods of low or no traffic (e.g., weekends), Cloud Run can scale the service down to zero instances, eliminating costs associated with idle GPUs and CPUs.
Model Loading Strategies
- Loading from the Public Internet on Demand: As demonstrated in the demo, the model can be loaded from the public internet when the
ollama runcommand is executed. The model remains in memory after loading. - Loading from Google Cloud Storage (GCS): The model can be downloaded to GCS, and the container can be configured to load it from there.
- Including the Model in the Container Image: The model can be included directly within the container image deployed to Cloud Run.
- Trade-offs: Each method has its own advantages and disadvantages, and the choice depends on factors such as startup time, network bandwidth, and storage costs.
Cloud Run GPUs: Key Advantages
- Fully Managed: Cloud Run GPUs are fully managed, eliminating the need for manual driver or library installations.
- On-Demand Usage: GPUs can be used on demand without requiring upfront reservations.
- Public Preview: Cloud Run GPUs are currently in public preview, requiring users to request enablement for their projects.
Conclusion
Cloud Run provides a streamlined and efficient platform for hosting and serving LLMs like Deepseek. By leveraging containerization, GPUs, and automatic scaling, developers can easily deploy AI-powered applications without the complexities of managing underlying infrastructure. The video demonstrates a three-step process for deploying Deepseek with Cloud Run GPUs, while also highlighting alternative approaches and best practices for model loading and scaling.
AI summaries can miss context or contain errors. Check important details against the original video.