How Can We Solve Observability's Data Capture and Spending Problem?

By The New Stack

Share:

Here's a comprehensive summary of the provided YouTube video transcript:

Key Concepts

  • Observability: The ability to understand the internal state of a system based on external outputs.
  • AI Workloads: Applications or components that leverage artificial intelligence, such as large language models (LLMs).
  • Golden Signals: Key metrics used to measure the health and performance of a system (latency, throughput, errors, saturation).
  • Accuracy as a Fifth Golden Signal: The emerging importance of evaluating the correctness and quality of AI-generated content.
  • Add-on Style AI Integration: Integrating AI as a supplementary service to an existing application.
  • Blocking Style AI Integration: Integrating AI as a core component within an application's workflow, where its failure impacts the entire application.
  • Hallucinations (AI): AI generating incorrect or fabricated information.
  • Prompt Engineering: The process of designing effective inputs (prompts) for AI models.
  • Telemetry Data: Data collected from various sources within an application and its infrastructure (metrics, logs, traces, events).
  • Per-Second Granularity: Collecting data at a very high frequency to capture transient events.
  • 100% Telemetry Capture: Collecting all available data, rather than sampling, to avoid missing critical information.
  • Cost of Observability: The significant expense associated with collecting, storing, and processing large volumes of telemetry data.
  • Business Leader Perspective: The need for observability tools to provide insights relevant to business outcomes, not just technical details.
  • Automated Prioritization: Observability platforms automatically ranking issues based on their impact on the business.
  • Gartner Magic Quadrant: A recognized industry report that evaluates vendors in specific technology markets.

State of Observability and Evolving Definitions

Jacob Jakonovich describes the state of observability as being in a constant state of change, with its definition and expected use cases evolving approximately 2% weekly. This evolution is driven by the rapid introduction of new technologies and principles, such as Kubernetes' enhanced support for AI workloads. The core challenge is to continuously redefine what observability tools need to do to understand why systems behave the way they do as these new elements are integrated into operations.

Observability for AI Workloads

The discussion highlights two primary approaches to integrating AI workloads:

  1. Add-on Style: AI is added as a supplementary service to an existing application. For example, a chatbot that allows users to query application data via natural language. In this model, if the AI component fails (e.g., provides a wrong answer), the core application functionality remains available, and the user can still complete their primary task.
  2. Blocking Style: AI is integrated as a critical part of the application's workflow. If the AI component generates incorrect, toxic, biased, or hallucinatory content, it can bring down the entire application or render it unusable. An example given is a flight booking system that returns a recipe for banana blueberry pie instead of flight information, leading the user to abandon the service.

The Emergence of Accuracy as a Fifth Golden Signal

With the rise of AI-generated content, traditional "golden signals" (latency, throughput, errors, saturation) are no longer sufficient. Accuracy is emerging as a crucial fifth golden signal. This involves evaluating the quality and correctness of the AI's responses from the end-user's perspective.

  • How it's Gauged: Observability platforms need to evaluate user prompts and AI responses. This includes checking for hallucinations, toxicity, bias, and overall quality.
  • Actionable Insights: If accuracy issues are detected, it signals a need for AI engineers and operations teams to retrain models, improve them, or even reconsider the choice of AI model.
  • Distinction from Traditional Observability: While metrics, logs, and traces remain critical for understanding the underlying infrastructure and microservice health, they don't inherently capture the accuracy of AI-generated content. Accuracy is about inspecting what goes into the AI, how the model processes it, and what comes back to the user.

Key Challenges in Observability

Several significant challenges are currently faced in the realm of observability:

  1. Pace of Business Change and User Interaction: Businesses are changing rapidly, and user interactions are becoming faster. This necessitates collecting data at a very high speed.
  2. Increasing Complexity of Applications: Applications are increasingly composed of a mix of legacy and cloud-native technologies, including numerous microservices. This creates a complex environment where components can have vastly different lifecycles, from long-lived to ephemeral.
  3. Missing Critical Data Due to Sampling or Polling: If data is not collected at per-second granularity and with 100% telemetry capture, transient issues, peaks, and valleys that indicate anomalies can be missed. This leads to a "white rabbit chase" for operations teams trying to diagnose problems without a solid foundation of information.
  4. Cost of Observability: The expense associated with collecting, ingesting, and storing vast amounts of telemetry data is a major barrier. Many organizations are forced to filter or limit data collection due to cost, potentially sacrificing comprehensive observability.
  5. Difficulty in Interpretation and Action: Even with collected data, it can be challenging for users, especially those without deep technical expertise, to parse through metrics, logs, and traces to identify problems and take appropriate actions.

Addressing the Cost of Observability

IBM's approach to observability emphasizes that cost should not force organizations to compromise on visibility. The goal is to provide comprehensive observability without making it prohibitively expensive. This involves:

  • Understanding Upfront Pricing: Customers should know the cost of observability before adopting new use cases or features, avoiding surprise bills that can amount to a significant percentage of annual recurring revenue.
  • Reanalysis of Technology and Partnerships: A reevaluation of the technology and partnerships required is necessary to achieve the desired level of granularity, intelligence, and understanding in operations at a reasonable cost.

Evolving Observability for Business Stakeholders

The trend is towards making observability more accessible and actionable for a wider audience, including business leaders.

  • Shifting from Technical Specificity to Business Impact: While skilled engineers can delve into flame graphs and line-of-code analysis, business leaders should not need this level of detail. They need insights that highlight what's important for the business.
  • Focus on Service Health: Business leaders think of their operations as services, not just individual applications or microservices. They want to know if their service is healthy and performing from the end-user's perspective.
  • Automated Prioritization: Observability tools should automatically prioritize issues and incidents based on their relative impact on the business. This allows operations teams to focus on the most critical problems.
  • Cultural and Process Integration: Achieving this requires a combination of cultural shifts, the right tools, and well-defined processes.

IBM's Recognition and Future Outlook

IBM has been recognized in Gartner's Leaders Quadrant for observability platforms, indicating their strong position in the market. The company is committed to continuing the journey of improving observability by enabling business leaders to define their operational perspectives and having observability tools automatically prioritize issues based on business impact.

Conclusion

The discussion underscores that observability is a dynamic field, increasingly influenced by the integration of AI. The emergence of accuracy as a critical metric, coupled with the persistent challenges of data volume, cost, and interpretability, necessitates innovative solutions. The future of observability lies in providing comprehensive, cost-effective, and business-centric insights that empower both technical teams and business leaders to understand and manage complex systems effectively.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video