Gemma 4 production stack: Model Armor, ADK Agents, Tracing
By Google Cloud Tech
Key Concepts
- Model Armor: A Google Cloud service used to detect malicious inputs (prompt injection, jailbreaking) and prevent sensitive data leaks (PII) at the network or application level.
- Load Balancer & URL Map: Used to route traffic to multiple backend services (VLM and Ollama) via a single endpoint, enabling efficient traffic management.
- Network Endpoint Group (NEG): A configuration object that represents a backend service (in this case, Cloud Run) for the load balancer.
- Service Extensions: Plugins that allow the load balancer to interact with external services like Model Armor to intercept and filter traffic before it reaches the model.
- ADK (Agent Development Kit): A model-agnostic framework for building AI agents, utilizing LiteLLM as a universal adapter to interface with various models (e.g., Gemma 4).
- Observability (Metrics vs. Tracing): Metrics (e.g., GPU utilization, token count) provide a high-level view of system health, while Tracing (e.g., Cloud Trace) provides granular, request-level debugging data.
- Sidecar Pattern: Deploying a secondary container (e.g., a Prometheus exporter) alongside the primary agent container to collect and export custom metrics.
1. Architecture and Traffic Management
The lab focuses on deploying Gemma 4 models using two different frameworks—VLM and Ollama—on Cloud Run. To manage these, a Regional Application Load Balancer is implemented.
- URL Mapping: A URL map allows the system to maintain a single endpoint. Requests are routed based on the path (e.g.,
/apifor Ollama,/v1for VLM). - Proxy-only Subnet: Reserved for the load balancer to communicate securely with private Cloud Run services within the VPC.
2. Security and Safety with Model Armor
Model Armor provides a robust defense layer that is more comprehensive than built-in model safety filters.
- Implementation: It is integrated at the network level via a Service Extension. This ensures that malicious prompts are blocked before they ever reach the backend model.
- Templates: Users define a
ModelArmorTemplatespecifying confidence thresholds for categories like hate speech, harassment, PII, and jailbreaking. - Default Response: If a threat is detected, the system returns a pre-configured "Guardian" error message, preventing the model from processing the malicious input.
3. Agent Development and Deployment
The "Guardian" agent is built using the ADK and LiteLLM.
- Local Testing: Developers use
ADK runorADK webto test agent logic locally before deployment. - CI/CD Pipeline: Cloud Build is used to automate the build and deployment process. The agent image is pushed to the Artifact Registry and deployed to Cloud Run.
- A2A (Agent-to-Agent) Protocol: Agents use an "Agent Card" to discover capabilities and communicate with other agents (e.g., the "Dungeon" monster).
4. Observability Framework
To manage production-scale agents, the presenters emphasize two pillars:
- Metrics: Using a Prometheus sidecar container, the system exports metrics like
generation_tokens_counterto Google Cloud Monitoring. This is critical for tracking costs and performance. - Tracing: Enabled via Cloud Trace (OpenTelemetry), this allows developers to inspect the latency of specific LLM calls and debug the agent's internal decision-making process.
5. Notable Quotes
- "Model armor is really versatile... you can use it in many different ways... it's detecting for malicious inputs as part of a prompt and also what it's going to be looking for sensitive data leaks." — Mayo
- "Metrics basically showing the 'what'... Tracing gives us more detail of what's going on behind the scenes... it helps you troubleshooting, analyzing, improve the overall AI system." — Annie
- "This is end-to-end agent system management... these aren't the responsibilities of a single individual. These are the responsibilities of a team." — Mayo
Synthesis and Conclusion
The episode demonstrates that moving an AI agent from a local prototype to a production-ready system requires a multi-layered approach. By combining Load Balancers for traffic control, Model Armor for network-level security, and Prometheus/Cloud Trace for observability, developers can create robust, scalable, and secure AI systems. The key takeaway is that production AI is a team effort, requiring coordination between data engineers, infrastructure specialists, and security experts to manage the complexities of modern agentic workflows.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

Is there a Chinese cyber threat to EU solar energy? | DW News
DW News

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial