Key Concepts
- Resource Optimization: Efficiently allocating and managing computing resources (CPU, memory) to applications.
- Kubernetes: A container orchestration platform for automating application deployment, scaling, and management.
- Phops (FinOps): A cloud financial management discipline focused on cost optimization and accountability.
- Requests: The minimum amount of resources (CPU, memory) a container is guaranteed to receive in Kubernetes.
- Limits: The maximum amount of resources a container is allowed to consume in Kubernetes.
- Cluster Autoscaler: A Kubernetes component that automatically adjusts the size of a cluster based on resource demands.
- Carpenter: An open-source node autoscaling tool for Kubernetes that optimizes node utilization and simplifies spot instance adoption.
- Spot Instances: Discounted, spare compute capacity offered by cloud providers like AWS, subject to interruption.
- On-Demand Instances: Standard, pay-as-you-go compute capacity offered by cloud providers.
- HPA (Horizontal Pod Autoscaler): A Kubernetes controller that automatically scales the number of pods in a deployment based on observed CPU utilization or other select metrics.
- Keda: Kubernetes Event-driven Autoscaling.
- Argo CD: A declarative, GitOps continuous delivery tool for Kubernetes.
- Flux: A GitOps operator for Kubernetes.
- CAdvisor: A container resource usage monitoring tool.
- KSM (kube-state-metrics): An add-on agent to generate and expose cluster-level metrics.
- Layered Configuration as Code: Managing configurations in layers, allowing for defaults and exceptions, and defining configurations as code.
- LLM (Large Language Model): A type of AI model trained on vast amounts of text data.
- Agentic AI: AI systems that can autonomously perform tasks and make decisions.
- Deterministic vs. Non-Deterministic: Deterministic systems produce the same output for a given input, while non-deterministic systems may produce different outputs.
Storm Forge Acquisition by CloudBolt
- Acquisition Details: Storm Forge, a Kubernetes resource optimization company, was acquired by CloudBolt, a phops platform. Yasmin Rajabi, former COO of Storm Forge, is now the Chief Strategy Officer at CloudBolt.
- Rationale: CloudBolt aims for continuous optimization of entire infrastructure. Storm Forge's Kubernetes-specific optimization complements CloudBolt's cost reporting and visibility capabilities. The goal is to apply Storm Forge's machine learning approach to CloudBolt's broader cloud infrastructure management.
User Challenges and Mistakes in Kubernetes Resource Management
- Initial Deployment Issues: Users often fail to set resource requests when initially deploying applications to Kubernetes. This leads to resource contention, performance issues, and potential downtime.
- Over-Provisioning: To avoid being paged, users often set resource requests too high, leading to wasted resources and increased costs.
- Manual Optimization: Users resort to manual resource adjustments, which are time-consuming and prone to errors.
- Scaling Challenges: As the number of clusters increases (dozens to hundreds), managing node count, node waste, and instance types becomes increasingly complex.
Solutions for Large-Scale Kubernetes Deployments
- Cluster Autoscaler: A starting point for autoscaling, but it may not be sufficient for advanced optimization.
- Carpenter: Helps right-size nodes, bin-pack workloads, reduce waste, and leverage spot instances.
- Machine Learning: Essential for handling the complexity of resource optimization at scale, considering factors like traffic patterns, scaling behavior (HPA), and custom metrics.
Machine Learning Approach
- Origin: Storm Forge started as a machine learning company focused on power optimization in data centers.
- Application to Kubernetes: The founding engineer recognized the potential of machine learning to solve resource optimization challenges in Kubernetes.
- Key Considerations: Machine learning algorithms must consider traffic patterns, scaling behavior, and custom metrics.
- User Experience: The user experience is crucial. Users need ways to constrain the machine learning and maintain control over the outcome.
- Configuration as Code: Users can interact with the machine learning via code, specifying constraints and considerations.
Information Architecture and User Experience
- Philosophy: Align with Kubernetes standards and best practices.
- Data Collection: Use standard Kubernetes metrics from tools like CAdvisor and KSM.
- Integration with Existing Tools: Work seamlessly with tools like HPA, Keda, Argo CD, Flux, and Carpenter.
- Kubernetes Language: Speak the Kubernetes language, so users don't have to learn new concepts.
Fitting into the CNCF Landscape
- Tooling Compatibility: Avoid forcing users to change their existing tools.
- Carpenter Integration: Work well with Carpenter for node autoscaling and spot instance adoption.
Technology Stack
- Proprietary Machine Learning: The machine learning algorithms are proprietary.
- Open Source Backend: The backend utilizes open-source solutions.
- Prometheus Agent: The agent that collects metrics from the customer's cluster is based on Prometheus.
Managing Exceptions at Scale
- Defaults and Exceptions: Allow users to set defaults for the entire estate but also manage exceptions for specific applications or workloads.
- Layered Configuration as Code: Use layered configuration as code to manage these exceptions.
- Boring Automation: The product should be "boring" and operate in the background, continuously right-sizing resources.
Machine Learning vs. AI
- Machine Learning as a Subset of AI: Machine learning is a subset of AI.
- Innovation vs. Fads: Distinguish between real innovation and fleeting fads in AI.
- Deterministic Requirements: Systems that touch production workloads must be deterministic.
- Non-Deterministic AI: AI systems, especially agentic AI, are often non-deterministic.
- Insights vs. Actions: Use AI for insights rather than automated actions to ensure safety and control.
CloudBolt Integration
- Cost Visibility: Layer on cost visibility into the platform, pulling data from billing and the curve.
- Broader Cloud Infrastructure: Apply the machine learning approach to the rest of CloudBolt's cloud infrastructure management capabilities beyond Kubernetes.
Conclusion
The acquisition of Storm Forge by CloudBolt aims to provide a comprehensive phops platform with continuous optimization capabilities. A key focus is on addressing the challenges of resource management in large-scale Kubernetes deployments through machine learning. The approach emphasizes user control, integration with existing tools, and a deterministic approach to automation. While AI offers potential for insights, the focus remains on reliable and predictable machine learning for production workload optimization.
AI summaries can miss context or contain errors. Check important details against the original video.





