Container-Optimized Compute Platform for GKE Autopilot

Google Cloud TechAbout 2 min readAug 7, 2025Watch original
THE SUMMARYAI-generated

Container Optimized Compute for GKE Autopilot: Faster Scaling

Key Concepts:

  • GKE Autopilot: Fully managed Kubernetes service.
  • Container Optimized Compute: Redesigned compute stack for faster scaling.
  • Horizontal Pod Autoscaler (HPA): Automatically scales the number of pods in a deployment.
  • In-place Pod Resize: Resizing pods without disruption (Kubernetes 1.33).
  • High Performance HPA Profile: Low-latency autoscaling with faster HPA calculation.
  • General Purpose Compute Class: Workload configuration for optimal performance with Container Optimized Compute.

1. The Challenge of Scaling in GKE Autopilot:

  • Historically, scaling in Kubernetes has been a challenge.
  • In GKE Autopilot, scaling requires creating new nodes before applications can scale onto them.
  • This node provisioning delay can be problematic for applications needing rapid scaling.
  • Users previously used "balloon pods" (dummy pods) to reserve nodes, which was costly and difficult to maintain.

2. Container Optimized Compute: A Solution for Faster Scaling:

  • Mission: Provide near real-time, vertically and horizontally scalable compute.
  • Goal: Deliver capacity when needed at the best price and performance.
  • Redesigned compute stack in GKE Autopilot.
  • Provides flexible compute on demand.
  • Results in up to 7x faster pod scheduling runtime.

3. Speeding Up Horizontal Pod Autoscaler (HPA):

  • Container Optimized Compute speeds up the HPA.
  • Leverages in-place pod resize (Kubernetes 1.33) for non-disruptive scaling.
  • All features are available out-of-the-box in GKE Autopilot.

4. High Performance HPA Profile:

  • Introduced for low-latency autoscaling reaction time.
  • Provides consistent horizontal scaling reaction time.
  • Offers up to 3x faster HPA calculation.
  • Uses higher resolution metrics for improved scheduling decisions.
  • Supports scaling up to 1000 HPA objects with predictive latency.

5. Demo of Container Optimized Compute:

  • Demonstrates rapid scaling by changing the replica count from 1 to 10.
  • Shows how quickly new pods are scheduled.

6. How to Use Container Optimized Compute:

  • Create a new GKE Autopilot cluster with the "HPA profile performance" enabled.
  • Ensure workloads use the "general purpose compute class."
  • Best suited for services that need to scale gradually.
  • Most improvement seen in workloads with smaller resource requests.

7. Limitations:

  • Not suitable for one pod per node deployments (anti-affinity situations).
  • Not suitable for batch workloads.

8. Conclusion:

  • Container Optimized Compute platform improves application autoscaling in GKE Autopilot.
  • Encourages users to try it out and provide feedback.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.