Improve GPU/TPU Obtainability with DWS Flex Start on GKE

Google Cloud TechAbout 4 min readMay 29, 2025Watch original
THE SUMMARYAI-generated

GKE Accelerator Obtainability Options: A Deep Dive

Key Concepts:

  • Accelerator Obtainability (GPUs, TPUs)
  • On-Demand vs. Spot Instances
  • Reservations (On-Demand, Future)
  • DWS (Delayed Workload Scheduling)
  • DWS Flex-Start (with discounted pricing)
  • Custom Compute Classes
  • Node Recycling
  • Queue (Kubernetes Batch Scheduling)
  • Gang Scheduling
  • Topology-Aware Scheduling

1. Challenges in Accelerator Usage

  • Hardware Demand: Continuous race for more performant hardware leads to high demand and limited availability. Hardware vendors struggle to meet the demand from AI companies.
  • Securing Access: Securing access to limited accelerator resources is crucial.
  • Cost Management: Balancing cost and performance is essential due to the significant expense of accelerators. Cloud bills can be substantial.

2. On-Demand vs. Spot Instances

  • On-Demand:
    • Provides full control over VM lifecycle.
    • Suitable for predictable workloads and no interruption.
    • Risk of hardware stockout.
  • Spot:
    • Offers great price flexibility.
    • Ideal for stateless workloads.
    • Prone to disruption.

3. Reservations

  • On-Demand Reservations:
    • Typical choice for mature AI companies needing large-scale infrastructure.
    • Consumed with discounting mechanisms (e.g., committed use discounts - CUDs).
    • Charged regardless of VM/node usage.
    • Suitable for large-scale critical applications like consistent model serving or large-scale training.
  • Future Reservations:
    • Allows defining the start and end date of the reservation.
    • Specify required capacity and virtual machine type.
    • Combined with DWS Calendar Mode to access newest accelerators (A3 Ultra, A4 VMs, upcoming types).
    • Dedicated recommender helps determine optimal dates based on requirements.
    • Ideal for platform evaluation, benchmarking, proof of concepts, and time-bound model training.

4. DWS (Delayed Workload Scheduling) Flex-Start

  • DWS Flex-Start: Schedules workloads and starts them as soon as possible in a delayed manner.
    • Improves accelerator obtainability with limited to no reservations.
    • Allows for gang scheduling (workload starts only when all required accelerators are available).
    • Relies on provisioning requests with queue.
  • Improvements to Flex-Start:
    • Discounted Pricing: Flex-Start is now a provisioning model balancing spot prices with on-demand availability.
    • Simplified Experience: As easy as using spot instances for smaller workloads (simple flag when setting up the notebook).
    • Integration with Compute Classes.
    • Upcoming support for TPUs in addition to GPUs.
  • Cost Optimization: Flex-Start sits between spot and on-demand in terms of cost.
  • Limitation: Nodes run for a maximum of seven days.
  • Node Recycling: Feature of custom compute classes to overcome the seven-day limit.
    • Automatically creates a new instance before the old one terminates.
    • Lead Time Parameter: Defines the period when the new node provisioning should start.
    • Crucial for model inference workloads requiring continuous operation.

5. Queue (Kubernetes Batch Scheduling)

  • Open-source batch scheduling system for Kubernetes.
  • Integrates with TFJobs, RayJobs, Kubernetes pods, and jobs.
  • Supports quota control, priorities with preemption.
  • Improves pod-to-pod communication with topology-aware scheduling.

6. Real-World Application Scenarios

  • Day 1 (AI Unicorn):
    • Requirements: 3 nodes with accelerators for platform benchmarking/PoC.
    • Single workload, no tight deadlines, 5 days of continuous computation, distributed training.
    • Recommendation: Calendar Mode or Flex-Start with queued provisioning (for gang scheduling).
  • Team Growth:
    • Requirements: Multiple teams, large workloads, up to 50 A3 machines.
    • Long-term reservation needed for long-running workloads to avoid disruptions.
    • Mixed workload (production and experimentation).
    • Solution: Long-term reservations for production, spot/Flex-Start for experimentation, queue for resource sharing, quota management for reserved instances.
  • Large-Scale Production System:
    • Requirements: Growing system (20% QoQ), product launch with unpredictable user traffic.
    • Solution:
      • Long-term reservations for baseline capacity.
      • Flexible capacity (Calendar Mode and Flex-Start) for growth.
      • Opportunistic capacity (spot instances) for marketing campaign traffic.

7. Choosing the Right Obtainability Option

  • Key Questions:
    • Workload type (inference, fine-tuning, mix)?
    • Specific hardware requirements (latest GPUs or more available ones)?
    • Budget (tradeoff between obtainability and cost)?
    • Criticality of immediate GPU access?
    • Need for guaranteed obtainability at a specific time?
    • Scalability requirements (rapid scale-up/down)?
  • Decision Matrix: Answering these questions leads to a matrix of choices to determine the best option.

8. Mapping Use Cases to Obtainability Options

  • DWS Calendar and Flex-Start are suitable for most use cases (except potentially giant model training).
  • Detailed breakdown of options includes discounting, SKUs, capacity guarantees, preemption behavior, quotas, and minimum commitment.

9. Conclusion

The video provides a comprehensive overview of accelerator obtainability options on GKE, emphasizing the importance of understanding workload requirements and balancing cost, availability, and performance. DWS Flex-Start and Calendar Mode offer flexible and cost-effective solutions for many AI/ML workloads, while reservations provide guaranteed capacity for critical applications. The integration with custom compute classes and queue simplifies the management and automation of these options.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.