This month in GKE: September edition

By Google Cloud Tech

Share:

Key Concepts

  • GKE Autopilot: A fully managed mode for Google Kubernetes Engine that automates infrastructure management.
  • Compute Class: A GKE feature enabling consistent, system-level configuration and tuning across nodes.
  • GKE Standard: The traditional GKE cluster mode where users manage nodes and infrastructure.
  • GKE Enterprise: A premium GKE offering with advanced features for large-scale, multi-cluster environments.
  • Fleet Dashboard: A centralized tool for managing and observing multiple GKE clusters.
  • Config Sync: A GitOps-based configuration management tool for GKE, ensuring cluster configurations are synchronized with a Git repository.
  • GitOps: An operational methodology that uses Git as the single source of truth for declarative infrastructure and application configurations.
  • OPA (Open Policy Agent): An open-source policy engine used for enforcing policies across various systems.
  • Policy Controller: An OPA-based tool within GKE for enforcing organizational policies within clusters.
  • L4 ILB (Layer 4 Internal Load Balancer): A load balancer operating at the transport layer, used for internal traffic within a Google Cloud network.
  • Zonal Affinity: A networking feature that prioritizes routing traffic to backends located in the same zone as the request origin.
  • Multi-subnet Clusters: GKE clusters that can utilize multiple subnets, allowing for greater IP address flexibility and scalability for pods and nodes.
  • Vertical Pod Autoscaler (VPA): A GKE feature that automatically adjusts CPU and memory requests/limits for pods based on their actual usage.
  • In-place Pod Resize: The capability to modify a pod's resource requests and limits without requiring the pod to restart.
  • Node Auto Provisioning (NAP): A GKE feature that automatically provisions new nodes to meet the demands of pending pods.
  • Kubernetes 1.34: The latest stable version of the Kubernetes container orchestration system.
  • Node Drain Timeout: The maximum duration allowed for a node to gracefully shut down and evict its running pods before being deprovisioned or updated.
  • A4X VM: A Google Cloud virtual machine type featuring Nvidia GB200 GPUs, specifically designed for high-performance AI and HPC workloads.
  • Nvidia GB200: A high-performance GPU from Nvidia, optimized for AI and high-performance computing.
  • GPU Fast Starting Notes: A feature that significantly reduces the startup time for nodes equipped with NVIDIA GPUs when using GKE Autopilot.
  • Service Agent: A Google-managed service account used by Google Cloud services to access resources on behalf of users.
  • CSI Driver (Container Storage Interface Driver): A standard interface that allows container orchestrators like Kubernetes to expose arbitrary block and file storage systems to containerized workloads.
  • Filestore Regional Tier: A regional file storage service in Google Cloud, offering managed network-attached storage (NAS).

GKE Autopilot and Standard Cluster Enhancements

  • Autopilot Compute Class Introduction: A new feature that extends the benefits of GKE Autopilot mode, such as faster pod scheduling and more efficient resource management, to standard GKE clusters. This allows users to leverage Autopilot's operational advantages without migrating to a new cluster, aiming for faster workload provisioning with minimal latency and maximum performance.
  • GKE Enterprise Features Consolidated into GKE Standard: Effective early September, several powerful features previously exclusive to GKE Enterprise are now available in GKE Standard. This consolidation makes advanced capabilities more accessible to all GKE users. These include:
    • Fleet Dashboard: For multi-cluster management and observability.
    • Config Sync: For GitOps-based configuration management, ensuring cluster configurations are managed declaratively from a Git repository.
    • OPA-based Policy Controller: To enforce policies within clusters, ensuring compliance and governance.
    • Multi-team Management Capabilities: For organizing and managing resources across different teams.

Networking Enhancements

  • Zonal Affinity for L4 ILB Services (Preview): This new preview feature prioritizes routing traffic for Layer 4 Internal Load Balancer (L4 ILB) services to backends located in the same zone as the request origin. If no healthy backend is available in the originating zone, the request automatically fails over to other regions, ensuring high availability.
  • Multi-subnet Clusters (Generally Available): The capability to add subnets to an already running GKE cluster is now generally available. This allows clusters to grow and accommodate more pods and nodes without the need to overprovision IP addresses, providing greater flexibility and scalability.

Autoscaling and Compute Class Improvements

  • Vertical Pod Autoscaler (VPA) with In-place Pod Resize (Preview): A new preview feature that allows the VPA to adjust the CPU and memory resources of workloads without disrupting or restarting the pods. This makes VPA a viable and highly beneficial solution for both stateless and stateful workloads, improving resource utilization and application stability.
  • Node Auto Provisioning (NAP) on a Per-Class Basis: Users can now enable Node Auto Provisioning (NAP) for specific compute classes, offering more granular control over how GKE provisions nodes.
  • Compute Class for System-level Configuration: Compute classes now allow users to define and apply system-level configurations and tuning across all their nodes in a consistent and maintainable way. This is a significant improvement over managing custom scripts or daemon sets for node-level configurations.

Kubernetes Version and Cluster Autoscaler Updates

  • Kubernetes 1.34 Availability: Kubernetes version 1.34 is now available on GKE, just two weeks after its open-source software (OSS) release, bringing the latest features and improvements from the Kubernetes community.
  • Increased Node Drain Timeout for GKE Cluster Autoscaler: The node drain timeout for the GKE cluster autoscaler has been increased from 10 minutes to one hour. This extended timeout provides more time for graceful shutdowns, which is particularly crucial for complex workloads such as AI and batch processing that require longer to complete their tasks before termination.

Hardware and Security Enhancements

  • A4X VM with Nvidia GB200 Availability: The A4X virtual machine, featuring the Nvidia GB200 GPU, is now available in GKE. The A4X VM provides a massive boost in performance for demanding AI and High-Performance Computing (HPC) workloads.
  • GPU Fast Starting Notes for Autopilot with NVIDIA GPUs: For users leveraging Autopilot with GPUs, the "fast starting notes" feature for NVIDIA GPUs significantly reduces the startup time of GPU-enabled nodes, improving the responsiveness of GPU-dependent workloads.
  • Dedicated Service Agent for Logging and Monitoring: A dedicated service agent for logging and monitoring GKE nodes is now available on GKE versions 1.33 and later, enhancing security and operational visibility.

Storage Options

  • Filestore Regional Tier Support by GKE CSI Driver: The GKE Container Storage Interface (CSI) driver now supports the Filestore Regional Tier. This provides yet another robust storage option for stateful applications running on GKE, offering regional availability and managed file storage capabilities.

Synthesis and Conclusion

This month's GKE updates emphasize a strong focus on simplification, enhanced performance, greater control, and increased accessibility for all users. By bringing GKE Enterprise features to GKE Standard and introducing the Autopilot Compute Class, Google is making advanced capabilities more broadly available and easier to manage. Significant improvements in networking, autoscaling (especially with VPA in-place resize), and hardware support (A4X VMs, GPU fast start) directly address performance and resource efficiency for diverse workloads, including demanding AI and HPC tasks. The extended node drain timeout and new security features further bolster the platform's reliability and operational robustness. Overall, these updates aim to streamline GKE operations, optimize resource utilization, and provide a more powerful and flexible environment for deploying and managing containerized applications.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video