THE SUMMARYAI-generated
Key Concepts
- Algorithm Audits: Systematic evaluations of algorithms to identify and mitigate bias, discrimination, and other unintended consequences.
- Sock Puppets: Fake online identities used to mimic real users in algorithmic audits, allowing researchers to observe how algorithms treat different demographic groups.
- Narrative Visualization Audits: Audits that involve interviews and visualizations to help people understand how algorithms curate their news feeds and other online experiences.
- Folk Theories: Non-authoritative conceptions of the world that develop among non-professionals and circulate informally, often used to explain how algorithms work.
- Usable Control: The ability for users to understand, access, and effectively use control settings to manage their online experiences and protect their autonomy.
- Contestability: The set of mechanisms for users to understand, construct, shape, and challenge model predictions.
- Placebo Effect: The phenomenon where people experience a benefit from a treatment or intervention, even if it is not actually effective.
- Expectations Hypothesis: The idea that people's expectations about how a control setting will work can influence their satisfaction with the setting, even if it is not actually effective.
- Infrastructure Calibrations: Studying the effects of changing one thing in large infrastructures to understand how they work and how they can be improved.
Saber System and the Origins of Algorithm Audits
- Saber System: Developed in 1953 and built in 1960 by American Airlines and IBM, it was a pioneering airline travel reservation system.
- Prior to Saber, reservations took about 90 minutes on average and were managed on paper.
- Saber managed the equivalent of 83,000 phone calls a day.
- In 1982, a study found that agents were choosing the top return selection on the Saber site more than half the time and flights on the first page more than 92% of the time.
- Manipulation Screen Science: Internally, the ranking algorithm design was intentionally called "manipulation screen science."
- American Airlines flights appeared first, even if they were more expensive or longer than competitors' flights.
- Antitrust Investigations: Complaints from competitors and travel agents led to antitrust investigations by the US Civil Aeronautics Board and the Department of Justice.
- Robert Crannle (CEO): Stated in 1983 that "The preferential display of our flights and the corresponding increase in our market share is the competitive reason for having created Saber in the first place."
- Regulation in 1984: The board decreed that the sorting algorithms for Saber must be known and passed a regulation requiring each airline reservation system to provide the current criteria used in editing and ordering flights, the weight given to each criterion, and the specifications used by the systems programmers in constructing the algorithm.
- Algorithm Audits Concept: Inspired by the Saber case and civil rights laws, the concept of algorithm audits was introduced in 2014.
Housing Audits and Sock Puppets
- Fair Housing Act of 1968 and Housing and Community Development Act of 1987: Inspired the work on housing audits.
- Traditional Housing Audits: HUD routinely performs housing audits using family pairs matched on everything except for one protected category.
- The first pair tests happened in 1977, and the most recent one found was in 2012.
- In the 2012 audit, blacks were shown about 17.7% fewer homes than whites, and Asians were shown about 18.8% fewer homes than whites.
- These audits take place roughly every decade and are very costly.
- Online Housing Discrimination: Questioned whether similar discrimination could be happening on websites that list homes.
- Sock Puppets: Unique fake identities online that directly mimic traditional audits.
- Thousands of agents were created, representing different demographic groups (white female, white male, black female, black male, etc.).
- Each agent starts as a fresh browser profile with no cookies or history.
- The profiles browse the internet following specific patterns, disproportionately visiting sites visited by members of their specific demographic group.
- Statistics on gender, age, income, ethnicity, and education were used to create these profiles.
- First Audit (Rankings): Women's feeds contained more expensive homes at the top compared to men's feeds.
- Second Audit (Advertisements): White people saw more housing ads than non-white people, and black people saw more ads for rent-to-own ads.
- New York Times Lawsuit (1991): The New York Times was sued for featuring thousands of white human models in their home section, while black models were depicted as maintenance employees, dormant small children, or cartoon characters.
Importance of Audits and Examples of Discrimination
- Facebook Advertising Discrimination: In October 2016, it was found that Facebook allowed advertisers to exclude users by race.
- Despite promises to stop this, ProPublica found in November 2017 that they were still excluding by race.
- In 2018, they promised to bar advertisers from targeting by race or ethnicity, but later that year, they were found of gender discrimination in job ads.
- In March 2019, they said they were going to do more to protect against discrimination in housing, but you could still discriminate against women and older workers despite a civil rights settlement.
- Obama Administration: Mentioned these audits as one of five national priorities essential for the development of big data technologies.
- They specifically said that we need to promote academic research and industry development of algorithmic auditing and external testing of big data systems to ensure that people are being treated fairly.
- Challenges: Audits require care, oversight, careful scrubbing of data, and sanity checks on data.
- In an arrest and recidivism audit, data was found to be unreliable due to the different data entry tools that were used and the interfaces that were used.
Narrative Visualization Audits and Folk Theories
- Narrative Visualization Audit: Led by Mojahari Islami, this audit included data with user permission and required an interview.
- Interested in seeing how people made sense of their news feeds and what approaches they used to get the content or the reception that they wanted.
- Awareness: Roughly 62% of people were not aware that there was an algorithm curating their feeds.
- People were blaming themselves for not seeing certain posts.
- Narrative Visualization Interface: Showed users the posts that were sent by everybody in their network and highlighted the ones that they saw.
- Showed users the posts that they did see and the people that see none of their posts, roughly half of their posts, and all of their posts.
- Reactions: Awareness and surprise were consistent, and several expressions around helplessness and loss of control emerged.
- Folk Theories: Non-authoritative conceptions of the world that develop among non-professionals and circulate informally.
- Personal Engagement Theory: Users thought that they could always click like on one of their own comments or status updates and this would trigger Facebook to show their post to more people.
- Control Panel Theory: Users felt that this gave them agency.
- Kemp's Research: Discovered a coexistence of folk theories and institutional theories in his work and found that each embraced theories each of these theories that people came up with produced both efficiencies and inefficiencies.
- Popularity: This type of audit became really popular and was used by thousands of people internationally.
Collective Audits and Community Engagement
- Collective Responses: Inspired many of the collective audits.
- Hotel Aggregator Systems: People were posting comments about how the algorithm didn't make any sense and were giving ideas.
- One aggregator site was rating their low to medium quality hotels 37% higher than other hotel rating platforms.
- People were reappropriating the website to make their own rating system.
- We Audit: A project led by Mutahari Islami getting people to come together to audit.
- Community Data Clinic: Working with local community organizations to help with community-centered audits.
- Multi-Disiplinarity: Requires a level of multi-disiplinarity that is not the norm.
- Sociotechnical Systems: Do not exist in a vacuum, and we are assessing systems with people that are not always rational.
- Theoretical Audits: Important, but more impact has been found when looking at measurement that responds to the real world with real individuals.
- Bins for People: Audits that only look at males and females are not accurate, and we should be looking at intersectionality.
Legal Challenges and Protections
- Lawsuit: Encountered challenges with the methods used to collect data, like web scraping and sock puppets, in courts.
- Federal prosecutors had interpreted the law to make it a crime to visit a website in a manner that violated terms of service.
- Filed a suit and found some protection for conducting audits.
- Federal court ruled that big data discrimination studies do not violate federal anti-hacking laws.
- The Department of Justice announced that they would not prosecute people who create fictional accounts for hiring housing or rental websites using pseudonyms on social networking sites that prohibit them or in violation of access restrictions in terms of service.
- AI Safety Research: Lagging, and present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness and barely address autonomous systems.
Control Settings and the Illusion of Control
- Control Settings: Most major platforms have a place where you can access a large control panel.
- Platforms love to tout how much control they're giving users.
- Navigation and Awareness Study: 50% of participants performed no better than chance at identifying which control settings existed.
- All participants struggled to find the control setting for at least one task.
- When guided through the settings, 94% of participants wanted to change at least one control setting.
- Users were misinterpreting settings, and their reasons for using settings did not quite match what the settings actually did.
- Placebo Study: Users were more satisfied with their news feeds when control settings were present, whether they were functioning or malfunctioning.
- When something didn't work the way they expected, people made excuses and blamed themselves.
- Financial Motives: There may be financial motives behind the design of control settings.
- We Metal: A tool built by Eric Gilbert in 2010 where users chose which algorithms they wanted to use in their feed.
- Replication Study: Replicated the placebo study at a larger scale using a probability sample across the US.
- Mechanisms: Expectations, sovereignty, clicking, and randomness.
- Results: The predictions of randomness, clicking, and the sovereignty hypothesis did not come true.
- Found a positive effect of control settings on feed selection, but the effect size was really small.
- Expectations Hypothesis: The bulk of the evidence happened through the path in the middle, which corresponds to the expectations hypothesis.
Choice vs. Automatic Systems
- Choice Study: Gave people the option of control by choosing when to see an ad on YouTube (pre-roll ad, mid-roll ad, or a choice between the two).
- Results: People prefer to get what they want automatically more than the choice over the ad placement.
- People expressed strong dislike for forced choice.
- Contestability: The set of mechanisms for users to understand, construct, shape, and challenge model predictions.
- Workshops: Organized workshops to discuss the potentials of and challenges designing for contestability.
- Marginalized Communities: Nine workshops were with members from three marginalized communities on Instagram that had their content removed.
- They wanted representation, to be valued by their peers, argumentation, cultural competence, compassion, and access.
Community-Based Work and Future Directions
- Community Levels: Looking at audits on online communities but also onlets in communities in our cities and our spaces.
- Center for Just Infrastructures and Community Data Clinic: Created a center and a community data clinic where people from the community can come and engage and create speculative audits.
- Local Communities: Have needs that don't always generalize.
- Infrastructure Calibrations: Looking at large infrastructures changing one thing and seeing what happens.
- Lee Star: We need a call to action to study boring things.
- Working with Lawyers: Being guided by law from step one.
- Content Moderation: Working to try to find audit mechanisms to protect free speech online while regulating content moderation.
Conclusion
- Well-constructed audits remain the best way to hold AI systems accountable.
- Combining control with audit can help us find a form of wanted autonomy.
- Usable control includes accurate messaging and contestability.
- Need to look at is it the sociotechnical system that is the problem or who has control over it and for what purpose.
- Need to replicate studies and monitor infrastructures.
- Need to look at them cognitively and sociotechnically and in and from many more perspectives.
AI summaries can miss context or contain errors. Check important details against the original video.