This month in GKE: August edition

Google Cloud TechAbout 3 min readSep 28, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Fast Starting GKE Nodes
  • Auto IPM (Automatic IP Address Management)
  • Multi-Subnet Clusters
  • GKE Inference Quickstart (GIQ)
  • Feature Gates for Alpha Clusters
  • Improved Horizontal Pod Autoscaler (HPA)
  • Secret Manager Add-on with Auto Rotation
  • Node Customization and Configurability (Max Image Pulls, Eviction Soft, Mswap, Topology Manager)
  • General Purpose (C4), Storage Optimized (Z3), and Memory Optimized (M4) Machine Types
  • Hyperdisk ML Volume Population

1. Fast Starting GKE Nodes:

  • Main Point: Significantly reduces cold startup times for GKE Autopilot nodes using Nvidia L4 Accelerators.
  • Specific Detail: New Nvidia L4 Accelerator can be initialized in under a minute.
  • Future Availability: Coming to GKE Standard soon, speeding up scale-up for services like Vertex AI.

2. TPU Startup Time Reduction:

  • Main Point: Reduced startup time by up to 50% for TPU v5e and v6e.
  • Actionable Insight: No code changes are required to benefit from this improvement.

3. Auto IPM (Automatic IP Address Management):

  • Status: Public Preview
  • Functionality: Dynamically allocates and manages IP addresses for GKE clusters.
  • Benefits: Simplifies IP planning, improves IP address efficiency, and allows clusters to scale without IP exhaustion.

4. Multi-Subnet Clusters:

  • Status: Public Preview
  • Functionality: Allows adding additional subnetworks to GKE clusters.
  • Use Case: Useful when a cluster is growing and needs more IP addresses for nodes or pods.
  • Benefit: Prevents the need to pre-allocate large IP address ranges, avoiding IP waste.

5. GKE Inference Quickstart (GIQ):

  • Status: Generally Available
  • Functionality: Analyzes, optimizes, and benchmarks AI models on GKE.
  • Output: Provides price-performance profiles and generates optimized manifests for model deployment.

6. Feature Gates for Alpha Clusters:

  • Functionality: Allows selectively enabling alpha and beta features for GKE alpha clusters.
  • Benefit: Enables focused testing without the instability of enabling all alpha features at once.

7. Improved Horizontal Pod Autoscaler (HPA):

  • Improvement: Able to handle up to 5,000 HPA objects per cluster.
  • Benefit: Faster reaction times for autoscaling events.

8. Secret Manager Add-on with Auto Rotation:

  • Enhancement: Secret Manager now supports Parameter Manager extension.
  • Key Feature: Supports auto-rotation for mounted secrets.
  • Functionality: Periodically updates and refreshes the content of secrets within pods when secrets are updated.

9. Node Customization and Configurability (Spencer's Segment):

  • Overarching Principle: Shifting configurability and customization to boot time for faster startup.
  • Max Image Pulls: Control the number of disk images pulled at once at boot to manage disk pressure.
  • Eviction Soft: Provides pods a grace period before being evicted due to CPU, memory, or disk pressure.
  • Mswap (Private Preview): Relieves memory pressure before OOMs and eviction soft kick in, improving P99 performance.
  • Topology Manager: Improves hardware placement for pods, reducing network latency and PCI Bus Express latency.

10. New Machine Types (Generally Available):

  • C4 (General Purpose): Intel Granite Rapids processors, local SSD storage.
  • Z3 (Storage Optimized): 3 TB to 72 TB of local SSD.
  • M4 (Memory Optimized): Up to 224 CPUs and 6 TB of memory, suitable for data workloads.

11. Hyperdisk ML Volume Population:

  • Status: Generally Available
  • Functionality: Automatically creates a Hyperdisk ML disk from a Cloud Storage bucket.
  • Benefit: Simplifies using large datasets for machine learning workloads.

Synthesis/Conclusion:

The August 2025 GKE update focuses on enhancing compute and networking capabilities. Key improvements include faster startup times for GKE nodes with accelerators and TPUs, better IP address management with Auto IPM and multi-subnet clusters, and tools for optimizing AI model deployment with GIQ. The update also provides more control over node configuration and introduces new machine types tailored for different workloads, alongside features like auto-rotating secrets and improved autoscaling. These updates aim to improve performance, scalability, and ease of use for GKE users.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.