Key Concepts
AI Agents, Language Models (LLMs), Short-term Memory, Long-term Memory, Context Sources, Task Orchestration, Synchronous vs. Asynchronous Agents, Human-in-the-Loop, Tools (Basic, Data Access, APIs, Image Generation, Browser, Code Execution), Cloud Run, Scalability, Cost Efficiency, Latency, Reliability, Developer Experience, Streaming, LangGraph, Agent Development Kit (ADK), Chains, Retrieval Augmented Generation (RAG), Tool Calling, State Machines, FastAPI, Docker, Google Cloud Build, Buildpacks, Artifact Registry, Vertex AI, Agent Builder, Agent Engine, Agent Garden, Model Garden, A2A Protocol, Model Context Protocol (MCP).
AI Agents: A New Application Architecture
AI agents represent a novel approach to application design, leveraging language models (LLMs) for reasoning and decision-making. Key characteristics include:
- Reasoning with LLMs: The core of an AI agent is its ability to use a language model to understand and process information.
- Memory: Agents possess both short-term and long-term memory to retain context and learn from interactions.
- Contextual Awareness: They connect to various data sources to gather relevant context for tasks.
- Task Orchestration: Agents can break down complex tasks into smaller, manageable steps, orchestrating multiple tools and even other agents.
- User Interaction: Agents can interact synchronously (e.g., via chat) or asynchronously (e.g., background research).
- Human-in-the-Loop: User moderation and control are crucial, allowing for guidance, confirmation, and decision-making.
Tools for AI Agents
Agents utilize various tools to perform tasks and interact with the external world:
- Basic Tools: LLMs are not good at math or basic time or unit conversions.
- Data Access: Agents can retrieve context from databases or other data sources (e.g., product catalogs).
- APIs: Agents can read from (and potentially write to, with supervision) first-party or third-party APIs.
- Image Generation: Agents can use image generation APIs or create visualizations from data.
- Browser: Agents can control a web browser to access and interact with websites.
- Code Execution: Agents can generate and execute code in a sandboxed environment.
Architecture of AI Agents
The architecture involves:
- User Interaction: Users interact with the agent through requests and streamed responses.
- Serving and Orchestration Component: This component uses a model for reasoning, memory, and access to data and tools.
Cloud Run as a Runtime for AI Agents
Cloud Run is presented as an ideal runtime environment for AI agents due to its:
- Scalability: Automatic, on-demand, and rapid scaling of instances.
- Cost Efficiency: Pay-per-use pricing with no flat fees or pre-provisioning.
- Managed Service: Fully managed with built-in zonal redundancy.
- Developer Experience: Easy deployment from source to URL with a single command.
- Security: Built-in credentials for secure access to the Gemini API.
- Flexibility: Support for any AI agent framework and language (Python, JavaScript, etc.).
- Streaming: Out-of-the-box HTTPS endpoints with streaming support (HTTP chunked transfer encoding, HTTP/2, WebSockets).
Implementation on Cloud Run
- Serving and Orchestration: A Cloud Run service handles serving and orchestration, running an AI agent framework (e.g., LangGraph, Agent Development Kit).
- Models: Access to the Gemini API, models served via Vertex AI, or fine-tuned models on Cloud Run with GPUs.
- Memory: Firestore (NoSQL database) or Memorystore.
- Retrieval Augmented Generation (RAG): Cloud SQL or AlloyDB for PostgreSQL with the pgvector plugin.
- Tools: Tools can be hosted on Cloud Run or accessed via existing APIs.
Building AI Agents with LangGraph on Cloud Run (Wietse Venema)
Wietse Venema demonstrates building an AI agent on Cloud Run using LangGraph.
LLMs and Their Limitations
LLMs are powerful but have limitations:
- Frozen Knowledge: Knowledge is limited to the time of training.
- Limited External Interaction: Cannot directly interact with the outside world.
Chains vs. Agents
- Chains: Fixed workflows with predetermined steps (e.g., Retrieval Augmented Generation - RAG). Advantage: Reliability. Disadvantage: Inflexibility.
- Agents: LLM in a loop, making decisions and calling tools. Advantage: Flexibility. Disadvantage: Lower reliability.
LangGraph: Combining Reliability and Flexibility
LangGraph uses a graph-based approach (state machine) to model control flow:
- Nodes: Represent specific decision points (code execution, LLM calls, tool calls).
- Edges: Define transitions between nodes, representing application logic.
LangGraph allows developers to explicitly define parts of the control flow, providing reliability where needed, while also allowing for the flexibility of agents.
Demo: Online Store Support Agent
The demo showcases an online store support agent built with LangGraph:
- Case Context Agent: Finds relevant resources related to a support case.
- SOP Agent: Reads a standard operating procedure (SOP), makes a plan, and executes it.
- Reply Agent: Sends a personalized message to the customer.
Application Architecture
- LangGraph: Used to build the agents.
- FastAPI: Web framework for the web application.
- Docker: Used to package the web app into a container image.
- Cloud Run: Deploys the container image and provides an HTTPS endpoint.
Deployment Process
The gcloud run deploy command:
- Uploads sources to a Google Cloud Storage (GCS) bucket.
- Kicks off a build in Cloud Build.
- Uses Buildpacks (if no Dockerfile is present) to create a container image.
- Pushes the container image to Artifact Registry.
- Cloud Run creates a new revision, importing the image.
- Automatically migrates 100% of traffic to the new revision.
Building AI Agents with Vertex AI (Sita Lakshmi Sangameswaran)
Vertex AI provides a comprehensive platform for the entire agent lifecycle.
Components of Vertex AI for Agent Building
- Models and Tools: Access to Google's Gemini models, open-source models via Model Garden, and custom models. Native Google Cloud integration for secure data access and tool connectivity.
- Vertex AI Agent Builder: Contains the Agent Development Kit (ADK), Agent Engine, and Agent Garden.
Agent Development Kit (ADK)
Google's open-source agent framework for building sophisticated and reliable agents.
- Built-in Developer UI: A local web interface for interacting with and debugging agents.
- Rich Interactions: Natively handles audio and video streaming.
- Open Ecosystem: Model agnostic, deployment agnostic, and interoperable. Supports Model Context Protocol (MCP) and Agent2Agent (A2A) Protocol.
Agent Engine
A serverless runtime optimized for agents. Handles scaling, security, and integration with cloud monitoring.
Agent Garden
Provides prebuilt agent samples and modular tool components.
Agent2Agent (A2A) Protocol
An open standard for enabling agents to securely communicate, exchange information, and coordinate actions. Supports various modalities (text, data, files) and enterprise-grade security. ADK natively supports A2A.
Synthesis/Conclusion
The presentation outlines two primary approaches to building AI agents on Google Cloud: leveraging Cloud Run for its scalability, cost-efficiency, and developer-friendly environment, and utilizing Vertex AI for its comprehensive platform, including the Agent Development Kit, Agent Engine, and Agent Garden. LangGraph is highlighted as a framework for building reliable and flexible agents by combining the strengths of chains and agents. The Agent2Agent (A2A) protocol is emphasized as a crucial element for enabling interoperability and collaboration between diverse agents, fostering a more interconnected and powerful AI ecosystem. Both Cloud Run and Vertex AI offer distinct advantages, catering to different development styles and deployment requirements, while the A2A protocol provides a unifying standard for agent communication and collaboration across platforms.
AI summaries can miss context or contain errors. Check important details against the original video.





