Key Concepts
Kubernetes 1.33, GKE Release Channels (Rapid), LLM Inference on Kubernetes (LLM-D), Performance Horizontal Pod Autoscaler (HPA), Confidential Nodes and GPUs, Container Optimized Compute (COC), Container Threat Detection (CTD), TPU in VLM, Network Analyzer for IP Masquerade, GKE AI Labs.
Kubernetes 1.33 Availability
- Kubernetes version 1.33 is available in the GKE Rapid channel.
- This release occurred two weeks after the open-source release, enabling early access to new features.
- Users can access it by selecting the Rapid release channel when creating or upgrading GKE clusters.
LLM-D: LLM Inference on Kubernetes
- LLM-D is a new open-source project aimed at simplifying large language model (LLM) inference at scale on Kubernetes.
- It's a collaborative effort between contributors to Kubernetes and the VLLM projects.
- LLM-D builds upon VLLM, Kubernetes, and the Inference Gateway.
- It provides a well-defined architecture for AI inference on Kubernetes.
Performance Horizontal Pod Autoscaler (HPA)
- A re-architected HPA stack is now available, offering significant performance improvements.
- It delivers two times faster autoscaling compared to the previous version.
- Improved metrics resolution enhances the accuracy of autoscaling decisions.
- The new HPA stack provides linear scalability, supporting up to 1000 HPA objects.
- It is the default on GKE Autopilot and available on standard mode clusters.
Confidential Nodes and GPUs
- GKE now supports confidential nodes and GPUs, enhancing data security.
- This feature leverages hardware-based encryption and integrity for data at use.
- GPU support includes the A3 H100, utilizing hardware encryption technology from the Nvidia Hopper Architecture.
- Standard VMs now support SEV-SNP (AMD) and TDX (Intel), hardware-based security standards.
- Support extends beyond N2D machines to include the Intel C3 architecture.
Container Optimized Compute (COC)
- Container Optimized Compute (COC) is now the default for GKE Autopilot clusters.
- This applies to both Rapid and Regular channels on Kubernetes version 1.32 or higher.
Container Threat Detection (CTD)
- Container Threat Detection (CTD) has been enhanced with support for new threats.
- New threat detections include Steganography Tool detection and command-line tool detection with known malicious patterns.
- Findings from CTD can be integrated into the Security Command Center for centralized security management.
TPU in VLM
- Support for Tensor Processing Units (TPUs) in VLM (Vision Language Model) is now generally available (GA).
- Official documentation is available on how to use TPUs in VLM within GKE.
Network Analyzer for IP Masquerade
- A new feature in Network Analyzer assesses and analyzes IP masquerade configurations for GKE nodes.
- It identifies potential issues with subnet ranges and network configurations.
- The tool helps users determine the correct subnetwork configurations to ensure proper pod masquerading and routing within the network.
GKE AI Labs
- GKE AI Labs is a new platform providing Kubernetes tutorials for AI and ML.
- It serves as a one-stop shop for resources and tools to facilitate AI adoption on GKE.
Conclusion
The May edition of "This Month in GKE" highlights several key updates focused on enhancing performance, security, and AI/ML capabilities within the GKE ecosystem. Key takeaways include the availability of Kubernetes 1.33 in the Rapid channel, the introduction of LLM-D for simplified LLM inference, performance improvements in the HPA, enhanced security with confidential nodes and GPUs, and new tools like the Network Analyzer and GKE AI Labs to streamline AI/ML development and deployment.
AI summaries can miss context or contain errors. Check important details against the original video.





