What's new for AI on GKE: Training, serving, and agents
By Google Cloud Tech
Key Concepts
- GKE (Google Kubernetes Engine): The managed Kubernetes service used for scaling AI workloads.
- Agentic Workloads: AI systems capable of autonomous reasoning, code execution, and task completion.
- gVisor: A sandboxing technology that provides VM-level security with container-like efficiency by intercepting syscalls in userspace.
- GKE Agent Sandbox: A service providing pre-provisioned, isolated environments for untrusted agentic code.
- POD24 Snapshots: A feature allowing the serialization of container/sandbox states to Google Cloud Storage to pause/resume idle workloads.
- Inference Gateway: A GKE feature for intelligent request routing to model replicas.
- Predictive Latency Boost: An ML-based routing optimization for Inference Gateway.
- LMDB (Large Model Distributed Backend): An open-source, CNCF-contributed framework for optimized distributed inference.
- Hairpinning: A networking technique where traffic between services in the same data center is routed internally rather than through external gateways to minimize latency.
1. Innovations in Agentic Workloads
Nathan Beach highlighted the shift from simple "one-shot" prompts (ChatGPT) to agentic systems that execute code and perform long-running tasks.
- Security: Because agentic workloads often execute untrusted code (e.g., LLM-generated scripts), gVisor is used to provide secure, isolated execution environments.
- GKE Agent Sandbox: Now generally available, it eliminates "cold starts" by maintaining a warm pool of sandboxes. It supports up to 300 new sandboxes per second per cluster.
- Cost Efficiency:
- Cold Standby Nodes: Allows maintaining a large pool of suspended sandboxes at a fraction of the cost.
- POD24 Snapshots: Enables "pausing" idle agents by serializing their memory and state to Cloud Storage, freeing up CPU/RAM for other tasks.
2. Inference Improvements
- Predictive Latency Boost: Replaces heuristic-based routing in the Inference Gateway with a continuously updated ML model, resulting in a ~70% reduction in "time to first token" latency.
- LMDB: A Kubernetes-native framework that handles complex tasks like KV cache offloading, tiered caching, and cross-host cache sharing.
- Asynchronous Inference: Allows GKE to prioritize real-time traffic while backfilling capacity with non-latency-sensitive batch jobs (e.g., reinforcement learning), maximizing hardware utilization.
3. Real-World Applications
- Base10: An inference provider that uses GKE to serve high-demand customers like Notion and Cursor. They utilize Dynamic Workload Scheduler to handle traffic bursts and hairpinning to achieve sub-millisecond latency between their inference engines and client clusters.
- Thinking Machines Lab: A research company that runs pre-training on Slurm (layered on GKE) and post-training/RL workloads. They emphasize Infrastructure as Code (IaC), which allows their internal agents to programmatically manage and modify their own infrastructure.
4. Methodologies for Efficiency
- Consolidation: Running diverse workloads (pre-training, inference, sandboxing) on a single platform (GKE) to simplify management.
- Fungible Compute: Striving for uniform compute environments (e.g., standardizing on A4/G4 instances) to allow workloads to be resized and mapped to hardware more flexibly.
- In-Cluster Caching: Deploying read-through caching layers directly on NVMe drives attached to GPU nodes to achieve terabytes-per-second throughput.
5. Notable Quotes
- “Software is eating the world... I think we can fairly say agents are eating the world.” — Nathan Beach
- “We don’t talk about nines of reliability. We talk about zeros. They want 100% or as close as we can get.” — Colin (Base10), regarding the critical nature of medical AI research.
Synthesis
The transition toward agentic AI requires infrastructure that is not only scalable but also secure and cost-efficient. Google’s strategy focuses on sandboxing (gVisor) for security, snapshotting (POD24) for resource optimization, and intelligent routing (Inference Gateway) for performance. By moving toward Kubernetes-native frameworks like LMDB and adopting IaC, organizations can automate the management of complex AI lifecycles, effectively turning infrastructure into a programmable asset that agents can optimize themselves.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
AI Engineer

'No where near normal' but 30-40 oil tankers passing through the Strait 'is better than 0': Mulberry
BNN Bloomberg

'Alphabet has such a dominant position they will be a leader in this space for many years': Clare
BNN Bloomberg

Forget Elon’s Data Centers In Space. This Startup Wants To Float Them At Sea
Forbes

Yahoo Finance Live: Daily Market Coverage - June 29, 2026 9AM-11AM (ET)
Yahoo Finance

Everyone's Buying AI. Smart Investors Are Buying This Instead. - Robert Kiyosaki
The Rich Dad Channel

Mad Money 06/26/26 | Audio Only
CNBC Television