Trustworthy Innovation for the Real-World in the Era of Foundation Models by Dr. Sirisha Rambatla

THE SUMMARYAI-generated

Trustworthy Innovation for the Real World in the Era of Foundation Models - Summary

Key Concepts:

  • Trustworthiness in AI: Reliability, coherence, alignment with expert knowledge, user trust, and national security.
  • Interpretability vs. Explainability: Inherently interpretable models vs. post-hoc explanations of black-box models.
  • Dictionary Learning: Decomposing data into a sparse combination of known patterns (prototypes).
  • Sparse Factor Models: Matrix factorization where the coefficient matrix is sparse.
  • Mechanistic Interpretability: Understanding the internal workings of a model, not just input-output relationships.
  • Alternating Optimization: Iteratively estimating dictionary and coefficients.
  • Low-Rank Adaptation: Fine-tuning large language models efficiently by exploiting the low-rank nature of gradients.
  • AI Literacy: User understanding of AI capabilities and limitations.

1. Introduction and Personal Journey

Dr. Sirisha Rrla, an assistant professor at the University of Waterloo, leads the Critical Machine Learning Lab. Her research bridges theoretical machine learning with real-world applications, focusing on reliability and trustworthiness. Her work spans AI for intelligent manufacturing (with Apple and Airbus/Naf Blue), AI for aviation operations (delay prediction), and AI for healthcare (liver transplantation, surgery, COVID-19 misinformation).

2. The Problem of Trust in AI: Early Examples

An early example of an app that could detect skin cancer from images was unreliable because tattoos could confuse the convolutional neural networks (CNNs). This highlighted the need to understand what CNNs are actually learning and what biases they might have. Research showed that CNNs could be biased towards texture rather than shape, raising questions about the reliability of their decisions. Early chatbots also exhibited biases and harmful behavior, emphasizing the need for control over what models learn.

3. Establishing Trust for Practitioners: Interpretability and Explainability

Interpretability is defined as the degree to which an observer can understand the causes of a prediction. ML models often capture correlations, not causations. Two approaches exist:

  • Interpretable Models: Inherently interpretable models like decision trees and linear regression. Decision trees are popular in healthcare because they can explain the reasoning behind decisions. Regularized linear regression promotes sparsity, making it easier to identify key features.
  • Black-Box Models: Explaining the decisions of black-box models (e.g., neural networks) post-hoc.

Anthropic's work on mechanistic interpretability aims to understand how large language models (LLMs) make decisions by analyzing individual neurons and decomposing their functions using dictionary learning.

4. Provable Algorithms for Dictionary Learning: Noodle

Dictionary learning involves representing data (Y) as a product of a dictionary matrix (A) and a sparse coefficient matrix (X). The challenge is that both A and X are unknown, and A can be overcomplete (wide), leading to more sparsity but also more challenges. The optimization problem is non-convex, making it difficult to guarantee finding the ground truth.

The main challenge is that any error in the estimated dictionary (A) propagates to the estimated coefficients (X), limiting guarantees. Existing methods using alternating optimization had an irreducible error term (order K/N, where K is sparsity and N is the length of dictionary elements), preventing exact recovery of the coefficients.

Dr. Rrla's work introduced a neurally plausible online dictionary learning algorithm (Noodle) that eliminates this irreducible error and achieves exact recovery of the sparse coefficient matrix. Noodle uses alternating optimization but avoids the error propagation issue.

Key Results:

  • Noodle achieves linear convergence to the global optimum.
  • It recovers both the dictionary and the coefficients accurately.
  • It is faster than state-of-the-art methods.

Noodle can be interpreted as a one-layer neural network with nonlinear activation, similar to the encoder used by Anthropic.

5. From Theory to Practice: Deep Fake Detection and Model Agnostic Explainability

While theoretical guarantees are valuable, they can be time-consuming to develop and rely on assumptions. Prototype-based models, where decisions are based on similarity to known prototypes (dictionary elements), offer a more practical approach.

Dr. Rrla's work used prototype-based models to explain deep fake detection, identifying specific features (e.g., eye movement) that led the model to classify a video as a deep fake.

For more complex models, explainability is crucial. A model-agnostic method called Archipelago was developed to identify feature interactions and their impact on model decisions. Archipelago breaks down detection from attribution, identifying which features are responsible and how they modify each other.

Examples:

  • Sentiment analysis: Archipelago correctly identifies that "bad," "terrible," "awful," and "horrible" are all negative and support each other, unlike previous methods.
  • Healthcare: Identifying comorbidities that, combined with age, lead to different outcomes.
  • Image classification: Identifying regions responsible for a particular output.

Archipelago helps practitioners trust models by showing that they are aligned with expert knowledge and not relying on spurious correlations (e.g., identifying birds based on the tree they are sitting on).

6. Trust for Users: Misinformation and the Limitations of LLMs

Generated images and LLMs can spread misinformation. Even if an LLM passes a medical licensing exam, users may rely on it for self-diagnosis, which can be risky.

A study tested GPT-4's correctness by presenting it with open-ended questions and then having experts and non-experts evaluate the responses. The success rate was only 31% when both groups agreed on the correctness. When information was dropped to simulate imperfect user input, non-experts categorized 27% more responses as correct than experts, highlighting the risk of spreading subtle misinformation.

7. Democratization of Training and Safety Alignment

LLMs are useful, but fine-tuning them requires significant computational resources, which is prohibitive for many organizations and contributes to climate change.

Dr. Rrla's lab is working on democratizing training by using low-rank structures to build more efficient fine-tuning methods. Recent research shows that even benign fine-tuning can degrade the safety alignment of LLMs, so it is crucial to ensure that fine-tuning is both democratized and safe.

8. Combating Misinformation and Ensuring Trust

To augment interpretability and explainability, it is essential to stop misinformation and contain harmful behavior. This is especially important as countries race to control these models.

Strategies include:

  • Real-time identification and stopping of misinformation.
  • Transparency about what is considered harmful.
  • AI literacy for users to understand the limitations of AI.

9. Conclusion

Establishing trust in AI requires a multi-faceted approach that includes theoretical guarantees, practical explainability methods, and strategies for combating misinformation. Democratizing training and ensuring safety alignment are crucial for making AI more accessible and reliable. Collaboration between experts from various fields is essential for addressing the complex challenges of trustworthiness in AI.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.