Deploying AI like it's code: A guide to upgrading agents

Google Cloud TechAbout 4 min readAug 29, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Agent architecture (agent, tools/services, model context protocol, agent-to-agent communication, APIs)
  • LLM (Large Language Model) upgrades and configuration
  • Prompt engineering ("prompts are code")
  • Containerization and runtime environments (e.g., Cloud Run)
  • Traffic splitting (0%, 1%, 10%, 50%, 100% traffic distribution)
  • A/B testing and monitoring (logs, performance)
  • RAG (Retrieval Augmented Generation) and vector databases
  • Embedding models and database migrations
  • API-driven service upgrades
  • Defense in depth
  • Software engineering principles for AI deployment

1. Agent Architecture and LLM Upgrades

  • The core of the system is an agent that interacts with various tools and services. These include:
    • A model context protocol for connecting to LLMs.
    • Other agents via agent-to-agent communication (ToA).
    • Traditional APIs for accessing services like order processing and product catalogs.
  • LLM upgrades are handled by treating the LLM as a configuration parameter within the agent.
  • Upgrading to a new LLM version often requires changes to system prompts, emphasizing the principle that "prompts are code" and should be version-controlled.

2. Deployment Strategy: Gradual Rollouts and Traffic Splitting

  • The recommended deployment strategy involves a gradual rollout using traffic splitting.
  • A new agent version (e.g., agent v2) is deployed with 0% traffic initially.
  • This allows for internal testing, diagnostics, and comparison against old logs to assess behavior before exposing it to users.
  • Traffic is then gradually increased (e.g., 1%, 10%, 50%, then 100%) while monitoring logs and performance metrics.
  • This approach minimizes risk and allows for early detection of issues.

3. Example: Pet Shop Application

  • The example used throughout the discussion is a "Pet Shop" application.
  • The agent interacts with services for order processing and accessing a product catalog.

4. RAG and Embedding Model Upgrades

  • The discussion extends to scenarios involving RAG, where the agent retrieves information from a vector database.
  • Upgrading the embedding model requires a database migration strategy.
  • A new column (e.g., "embedding V2") is added to the database to store embeddings generated by the new model, allowing both old and new embeddings to coexist.
  • Alternatively, a new database with different IDs can be created.

5. API-Driven Service Upgrades

  • The product catalog is treated as a separate service accessed via an API.
  • This allows for independent upgrades of the catalog service, including the use of the new embedding model for search.
  • Agent v2 can connect to catalog v2, using the new embedding model to generate embeddings for queries and lookups.
  • This approach provides "defense in depth," enabling incremental upgrades of system components without requiring a complete system shutdown or a risky "yolo" deployment.

6. Software Engineering Principles

  • The presenters emphasize that the deployment strategies discussed are based on standard software engineering principles.
  • These include separation of concerns, scaling out, and monitoring.
  • The key difference in AI deployments is the need for more robust evaluation methods to ensure that the non-deterministic AI system meets user goals.

7. Notable Quotes

  • "Prompts are code" - Emphasizing the importance of version control and testing for prompts.
  • "It's just software engineering" - Highlighting the applicability of traditional software development practices to AI deployments.

8. Technical Terms

  • LLM (Large Language Model): A powerful AI model used for natural language processing.
  • RAG (Retrieval Augmented Generation): An approach that combines information retrieval with text generation.
  • Embedding: A vector representation of text or other data used for semantic search and similarity comparisons.
  • Vector Database: A database optimized for storing and querying vector embeddings.
  • Traffic Splitting: A technique for gradually routing traffic to a new version of a service.
  • Containerization: Packaging an application and its dependencies into a container for consistent deployment.
  • Runtime: The environment in which a containerized application runs (e.g., Cloud Run).

9. Logical Connections

  • The discussion starts with basic agent architecture and then moves to LLM upgrades.
  • It then introduces deployment strategies using traffic splitting and monitoring.
  • The conversation extends to RAG and embedding model upgrades, highlighting the need for database migrations and API-driven service upgrades.
  • The presenters emphasize that all these concepts are grounded in standard software engineering principles.

10. Conclusion

The key takeaway is that deploying AI-powered applications, while having unique challenges like prompt engineering and non-deterministic behavior, can leverage established software engineering practices. Gradual rollouts, traffic splitting, and API-driven service upgrades are crucial for minimizing risk and ensuring a smooth transition when upgrading LLMs, embedding models, or other components of the system. The presenters stress the importance of treating prompts as code and continuously monitoring performance to ensure that the AI system meets user goals.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.