Understanding and Addressing Fairwashing in Machine Learning by Sébastien Gambs

THE SUMMARYAI-generated

Key Concepts

  • Fair Washing: The practice of manipulating explanations or communication to create the perception that a machine learning model is fair, when in reality it is not.
  • Ethics Washing: A broader term encompassing fair washing, greenwashing, and privacy washing, where an organization or individual falsely claims to adhere to ethical considerations to improve their public image.
  • Greenwashing: A company falsely claiming to be environmentally friendly.
  • Privacy Washing: A company falsely claiming to protect user privacy.
  • Big Data: The increasing volume, variety, and velocity of data, often personal data, used in machine learning.
  • Machine Learning (ML): Algorithms that enable systems to learn from data and make predictions or decisions.
  • Bias in Data: Societal and historical biases present in data that can be replicated and amplified by ML models.
  • Predictive Justice: Using ML models to assess the risk of recidivism for individuals, impacting decisions on release.
  • GDPR (General Data Protection Regulation): EU regulation mandating fairness and explainability in data processing.
  • AI Act (EU): Proposed EU regulation for AI systems, with specific requirements for high-risk AI.
  • Individual Fairness: The principle that similar individuals should receive similar predictions.
  • Group Fairness: The principle of equalizing statistical outcomes across different demographic groups.
  • Explainability (XAI): Techniques used to understand why an ML model makes a particular prediction, often applied post-hoc.
  • Interpretability: Designing ML models that are inherently understandable by humans (e.g., decision trees, rule lists).
  • Differential Privacy: A privacy model that provides strong privacy guarantees by adding noise to data.
  • Epsilon (ε): A privacy parameter in differential privacy that controls the level of noise and protection.
  • Rashomon Effect (Predictive Multiplicity): The phenomenon where multiple models can achieve optimal accuracy, creating a space of possible explanations and models.
  • Zero-Knowledge Proofs (ZKPs): Cryptographic methods that allow one party to prove to another that a statement is true, without revealing any information beyond the truth of the statement itself.
  • ZKML (Zero-Knowledge Machine Learning): The application of ZKPs to prove properties of ML models.

Fair Washing: Manipulating Explanations to Conceal Unfairness

This presentation introduces the concept of "fair washing," a form of "ethics washing" where the perceived fairness of machine learning (ML) models is manipulated, particularly through the explanation of their decisions. This practice is becoming increasingly relevant due to growing regulatory pressure for fairness, privacy, and explainability in AI systems.

The Rise of Big Data and Machine Learning

The past 15 years have seen a significant increase in the availability and use of "big data," much of which is personal data. This, coupled with advancements in ML, has led to the integration of ML models into various aspects of daily life, impacting humans directly. Examples include:

  • Medical Domain: ML models assisting doctors in identifying tumors in X-rays.
  • Predictive Justice (USA): ML models used to assess recidivism risk, influencing decisions on releasing individuals.
  • Automated Trading: ML models driving financial market transactions.

Regulatory Landscape and the Need for Fairness

The impact of ML models on individuals has prompted regulatory action. The EU's AI Act, for instance, identifies high-risk domains such as critical infrastructure, education, employment, law enforcement, and justice, where ML integration requires careful consideration of fairness, privacy, and security.

Historically, ethical considerations were addressed through declarations like the Montreal Declaration for Responsible Development of AI. However, these have had limited practical impact as they are not legally binding. Consequently, there's a shift towards regulatory enforcement.

Regulations like the GDPR in Europe already mandate explainability and fairness. Even before the AI Act, GDPR stipulated a right to an explanation for decisions made by AI systems and required data controllers to implement mathematical and statistical procedures to minimize discrimination risks based on sensitive attributes.

The Problem of Data Bias

A fundamental issue is that data often reflects societal and historical biases. Blindly training ML models on such data leads to the replication and amplification of these biases. A notable example is the COMPAS tool used in predictive justice, where a study found that African Americans were twice as likely to be predicted as high-risk for recidivism compared to white Caucasians. Similarly, in salary prediction, historical gender pay gaps can lead ML models to predict lower salaries for women.

Ethics Washing: A Growing Concern

Ethics washing, a concept studied in marketing for a long time, is the practice of feigning ethical considerations to improve perception. Greenwashing (falsely claiming environmental friendliness) and privacy washing (falsely claiming privacy protection) are examples.

Privacy Washing Examples:

  • A company merely repeating privacy regulations on its website without implementing more robust measures.
  • Promoting the use of privacy-enhancing technologies (PETs) without disclosing full details, leading to potential privacy leakage.
  • Using PETs as a justification for collecting more data or to hinder auditing by claiming data is encrypted or secret-shared.

Fairness in Machine Learning: Definitions and Challenges

Fairness in ML is a complex and actively researched area. Unlike privacy, where differential privacy is often considered a gold standard, there are numerous ways to define fairness.

  • Individual Fairness: Requires that similar individuals receive similar predictions. If two CVs are identical except for gender, individual fairness implies the same prediction.
  • Group Fairness: Focuses on statistical parity across groups, such as equalizing the probability of recruitment for men and women.

A key challenge is that these two notions of fairness can be mutually exclusive. Enforcing group fairness (e.g., equalizing acceptance rates) might break individual fairness, potentially leading to positive discrimination where individuals from a disadvantaged group with the same profile receive a higher score.

Methods to Improve Fairness:

  • Data Preprocessing: Modifying data to remove biases (e.g., equalizing salaries for individuals with similar profiles).
  • Fairness-Aware Training: Integrating fairness metrics into the ML model's optimization objective alongside accuracy.
  • Post-processing: Adjusting model weights or outputs to reduce discrimination.

It's crucial to note that achieving perfect fairness might come at the cost of utility. A coin flip, while perfectly fair, offers no predictive power.

Fair Washing in Practice: The COMPAS Case Study

Fair washing aims to make an ML model appear fair, especially when audited, even if fairness was not a primary consideration during its initial development. The COMPAS tool controversy illustrates this:

  • A report by ProPublica highlighted the unfairness of COMPAS, showing disparate recidivism risk predictions for African Americans.
  • The company behind COMPAS (now Northpointe, then Equivant) responded by arguing that the model was fair according to specific group fairness metrics, suggesting the original accusation was based on a misinterpretation of fairness.

This highlights how companies can leverage the multitude of fairness definitions to select one that makes their model appear fair, even if other, more critical, fairness metrics are violated.

Explainability and Interpretability

  • Interpretability: Using models that are inherently understandable by design (e.g., decision trees, rule lists, simple scoring systems).
  • Explainability (XAI): Techniques applied post-hoc to understand the predictions of complex "black-box" models (e.g., deep neural networks). This can involve:
    • Global Explanations: Training an interpretable model to mimic the black-box model.
    • Local Explanations: Techniques like LIME (Local Interpretable Model-agnostic Explanations) that highlight the input features most responsible for a specific prediction. For example, LIME might show that sneezing and headache were key predictors of the flu, while lack of fatigue was not.

Manipulating Explanations: The Foundation of Fair Washing

It is known that explanations themselves can be manipulated. Research has shown that saliency maps, used to highlight important pixels in an image for a prediction, can be adversarially modified without human perception of difference, leading to misleading explanations.

LaundryML: An Algorithm for Fair Washing

A research effort introduced "LaundryML," an algorithm designed to perform fair washing by manipulating explanations. The goal is to decrease the perceived unfairness of a model when presented to an auditor, even if the original model was unfair.

Methodology:

  1. Model and Outcome Rationalization: Manipulating explanations to hide original unfairness.
  2. Fairness Metric: Demographic parity (difference in positive outcome probability between groups).
  3. Explanation Quality Metric: Fidelity (agreement between the interpretable explanation model and the black-box model).
  4. Building Blocks: Supervised ML models (e.g., Corell) and model enumeration techniques to explore the space of interpretable models.
  5. LaundryML Algorithm: A modified version of Corell that integrates unfairness into the search space, aiming to find interpretable models that are small, accurate, and less unfair than the original black-box model.

Findings:

  • Experiments on datasets like Adult and COMPAS using Random Forests as black-box models demonstrated that LaundryML could produce Pareto fronts showing a trade-off between fidelity and unfairness.
  • This means that by choosing a specific fidelity threshold (e.g., 90% faithfulness), an adversary could select an interpretable model that appears significantly less unfair than the original model.
  • Feature importance analysis showed that after fair washing, features like gender, which were important in the original model, became less prominent in the explanation.

Why is Fair Washing Possible? The Rashomon Effect

The Rashomon Effect, or predictive multiplicity, explains why fair washing is possible. It suggests that for many ML tasks, there isn't a single "best" model. Instead, there exists a space of models with optimal accuracy. Similarly, there can be multiple explanations that are faithful to a black-box model. An adversary can exploit this by exploring this space and choosing an explanation that appears fair, even if the underlying model is not.

Fair washing is not limited to specific model types; it has been observed across logistic regression, AdaBoost, deep neural networks, and Random Forests, with varying degrees of effectiveness.

Detecting Fair Washing

Detecting fair washing is challenging. While some statistical group analysis can help, it is often fragile and can be circumvented. Accessing only an API makes detection even more difficult.

The Role of Regulation and Standards

Fair washing thrives in the absence of specific regulations and standards. The GDPR's broad definition of "explanation" allows for interpretation, and the lack of standardized fairness metrics enables companies to select metrics that favor their models. Proving intentionality behind fair washing is also extremely difficult.

Towards Verifiable Fairness with Cryptography

To combat fair washing, there's a growing interest in using cryptography to prove properties of ML models without revealing sensitive data or the model itself. This is often referred to as Zero-Knowledge Machine Learning (ZKML).

  • Proposed Solution: Establishing precise standards for fairness metrics and explanation techniques for specific domains (e.g., finance, insurance).
  • Cryptographic Proofs: Using techniques like Zero-Knowledge Proofs (ZKPs) to demonstrate that a model was trained according to these standards.
  • Examples:
    • Building fair decision trees and proving their fairness using ZKPs.
    • Research on individual fairness for neural networks.
    • End-to-end ZKPs for fairness from preprocessing to the final model.

This approach allows auditors to verify compliance without accessing proprietary data or models, making it harder to engage in fair washing.

Potential for Cryptography in Legal Defenses

In the US, legal provisions may allow for discrimination if a company can prove no highly accurate alternative model exists. The Rashomon effect, when combined with cryptographic proofs, could potentially be used to demonstrate that the chosen model is the most accurate and least unfair among all models in the "Rashomon set," thus supporting a legal defense.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.