Gemma 4 production stack: Model Armor, ADK Agents, Tracing

By Google Cloud Tech

Share:

Key Concepts

  • Model Armor: A Google Cloud service used to detect malicious inputs (prompt injection, jailbreaking) and prevent sensitive data leaks (PII) at the network or application level.
  • Load Balancer & URL Map: Used to route traffic to multiple backend services (VLM and Ollama) via a single endpoint, enabling efficient traffic management.
  • Network Endpoint Group (NEG): A configuration object that represents a backend service (in this case, Cloud Run) for the load balancer.
  • Service Extensions: Plugins that allow the load balancer to interact with external services like Model Armor to intercept and filter traffic before it reaches the model.
  • ADK (Agent Development Kit): A model-agnostic framework for building AI agents, utilizing LiteLLM as a universal adapter to interface with various models (e.g., Gemma 4).
  • Observability (Metrics vs. Tracing): Metrics (e.g., GPU utilization, token count) provide a high-level view of system health, while Tracing (e.g., Cloud Trace) provides granular, request-level debugging data.
  • Sidecar Pattern: Deploying a secondary container (e.g., a Prometheus exporter) alongside the primary agent container to collect and export custom metrics.

1. Architecture and Traffic Management

The lab focuses on deploying Gemma 4 models using two different frameworks—VLM and Ollama—on Cloud Run. To manage these, a Regional Application Load Balancer is implemented.

  • URL Mapping: A URL map allows the system to maintain a single endpoint. Requests are routed based on the path (e.g., /api for Ollama, /v1 for VLM).
  • Proxy-only Subnet: Reserved for the load balancer to communicate securely with private Cloud Run services within the VPC.

2. Security and Safety with Model Armor

Model Armor provides a robust defense layer that is more comprehensive than built-in model safety filters.

  • Implementation: It is integrated at the network level via a Service Extension. This ensures that malicious prompts are blocked before they ever reach the backend model.
  • Templates: Users define a ModelArmorTemplate specifying confidence thresholds for categories like hate speech, harassment, PII, and jailbreaking.
  • Default Response: If a threat is detected, the system returns a pre-configured "Guardian" error message, preventing the model from processing the malicious input.

3. Agent Development and Deployment

The "Guardian" agent is built using the ADK and LiteLLM.

  • Local Testing: Developers use ADK run or ADK web to test agent logic locally before deployment.
  • CI/CD Pipeline: Cloud Build is used to automate the build and deployment process. The agent image is pushed to the Artifact Registry and deployed to Cloud Run.
  • A2A (Agent-to-Agent) Protocol: Agents use an "Agent Card" to discover capabilities and communicate with other agents (e.g., the "Dungeon" monster).

4. Observability Framework

To manage production-scale agents, the presenters emphasize two pillars:

  • Metrics: Using a Prometheus sidecar container, the system exports metrics like generation_tokens_counter to Google Cloud Monitoring. This is critical for tracking costs and performance.
  • Tracing: Enabled via Cloud Trace (OpenTelemetry), this allows developers to inspect the latency of specific LLM calls and debug the agent's internal decision-making process.

5. Notable Quotes

  • "Model armor is really versatile... you can use it in many different ways... it's detecting for malicious inputs as part of a prompt and also what it's going to be looking for sensitive data leaks." — Mayo
  • "Metrics basically showing the 'what'... Tracing gives us more detail of what's going on behind the scenes... it helps you troubleshooting, analyzing, improve the overall AI system." — Annie
  • "This is end-to-end agent system management... these aren't the responsibilities of a single individual. These are the responsibilities of a team." — Mayo

Synthesis and Conclusion

The episode demonstrates that moving an AI agent from a local prototype to a production-ready system requires a multi-layered approach. By combining Load Balancers for traffic control, Model Armor for network-level security, and Prometheus/Cloud Trace for observability, developers can create robust, scalable, and secure AI systems. The key takeaway is that production AI is a team effort, requiring coordination between data engineers, infrastructure specialists, and security experts to manage the complexities of modern agentic workflows.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video