Stanford CS329H: Machine Learning from Human Preferences | Autumn 2024 | Introduction

Unknown AuthorAbout 5 min readSep 14, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Trustworthy Machine Learning and AI
  • Machine Learning from Human Preferences (MLHP)
  • Interactive Querying
  • Elicitation of Values and Preferences
  • Query Efficiency
  • Human-in-the-Loop Systems
  • Human Computer Interaction (HCI)
  • Reward Hacking
  • Reinforcement Learning from Human Feedback (RLHF)
  • Direct Preference Optimization (DPO)
  • Metric Elicitation
  • Inverse Decision Theory
  • Choice Models

1. Course Overview and Framing:

  • Course Goal: To explore the challenges of efficiently and effectively eliciting values and preferences from individuals, groups, or societies and embedding them within AI models and applications.
  • Focus: Statistical and conceptual foundations and strategies for interactively querying humans to elicit information that can improve learning and applications.
  • Timeliness: The course is timely due to the increasing engagement in machine learning aspects that explicitly use human feedback.
  • New Course: This is the second time the course is being offered, with improvements based on feedback from the previous iteration.
  • Textbook: A textbook specifically for this topic has been created by the core staff.
  • Course Structure: The course is divided into an introduction and four modules: modeling human choice, model-based techniques and preference learning, model-free optimization, and human values and AI alignment.
  • Assessment: 60% homework, with the first homework due in about a week, and 40% project and class participation.
  • Projects: Students will organize into groups of up to five for a final project, presented to the class and potentially external audiences.
  • Class Participation: Encouraged through in-class discussions, feedback on the textbook on GitHub, and engagement on ED.
  • CSPD Students: The course counts towards either learning and modeling requirements or human and society requirements.
  • Technical Course: The course is technical, focusing on machine learning fundamentals, but also considers human engagement and societal considerations.

2. Interactivity and Human Signal:

  • Interactive Focus: The course focuses on settings where the human signal is explicit and intentional, often through interactive processes.
  • Implicit vs. Explicit Preferences: Distinguishes between implicit preferences (e.g., learning from labeled data) and explicit preferences (e.g., interactive querying).
  • Human Inconsistency: Addresses the inconsistency of human labels and how to model the labeling process.

3. Foundations and Applications:

  • Foundations: Draws from economics, psychology, marketing, and statistics.
  • Applications: Includes language models, robotics, and logistics.
  • Machine Learning Lens: Focuses on modeling, estimation, and evaluation within each problem setting.
  • Prerequisites: Assumes comfort with machine learning basics, including train/test/validation splits and logistic regression.

4. Human Preferences and Societal Considerations:

  • Human Aversion: Notes the aversion to thinking about humans in the machine learning world, despite their pervasive influence.
  • AI as HCI: Presents the argument that within a generation, all of AI will be HCI, with the hardest questions being about human-computer interaction.
  • Sub-questions: Explores biases, rationality assumptions, human errors as noise, correctness, expertise, and human engagement with AI systems.
  • Individual vs. Group Preferences: Considers the differences between eliciting preferences from individuals versus groups or societies.
  • Ethical Considerations: Addresses ethical angles, such as the impact of choosing whose preferences to align with and potential exploitation.

5. Sampling and Representation:

  • Human Sampling: Discusses the need for human sampling to ensure consideration of different groups and ethical human preferences.
  • Limited Samples: Acknowledges the challenge of learning with limited samples due to the expense of querying humans.

6. Project Scope:

  • Project Types: Theoretical research, critical examination, literature surveys, and technical research with code are all within scope.
  • Application Areas: Language, vision, legal, and policy applications are all relevant.

7. Scope and Limitations:

  • Emerging Topic: The topic is emerging, with unclear edges and broad scope.
  • Breath Bias: The course has a breath bias, covering many topics without being exhaustive.
  • HCI Depth: The course will not delve deeply into HCI problems, but experience in HCI is beneficial.

8. Human Feedback in Machine Learning:

  • Human Feedback Shapes ML: Human feedback shapes data selection, labeling, model selection, training, evaluation, deployment, and context.
  • Taxonomizing Human Preferences: There is work taxonomizing ways that human preferences shape machine learning models.

9. Examples and Applications:

  • GPT and RLHF: GPT's success is attributed to aligning with human preferences via RLHF.
  • Language Model Feedback: Includes document-level interaction, label consistency queries, word-level interaction, feature selection, and model parameter selection.
  • Parsing Improvement: Improving parsers by labeling connections between words.
  • Supervised Fine-Tuning: Using knowledgeable humans to provide examples of questions and answers.
  • Comparison Data: Querying humans to get preference signals between different kinds of completions for a certain query.
  • Pairwise Preferences: Asking humans to choose between completions.
  • Optimization: Using reinforcement learning strategies like PPO to get models to pick completions that humans would prefer.

10. Reward Models and DPO:

  • Reward Model Training: Training a reward model separately and then training a language model to fit the reward model tends not to work well.
  • DPO: Direct Preference Optimization does away with explicitly learning the reward model.
  • Reward Model Value: The reward model itself is valuable as an artifact because it hopefully tells us something about human preferences.

11. Why Human Preferences Matter:

  • No Explicit Metric: Useful when there is no good explicit metric or loss function.
  • Stakeholder Care: Important when stakeholders care a lot about the outcome.
  • Ideal Behavior: Useful when there is some ideal behavior, but we don't know how to fully specify it.
  • Proxy for Real-World Outcome: A useful proxy for some real-world outcome that we'd like to get.

12. Challenges and Biases:

  • Explicit Biases: Language models have explicit biases that reflect human biases.
  • Unreliable Preferences: Human preferences can be unreliable.
  • Ethical Issues: Potential ethical issues related to outsourcing data labeling to low-cost settings.
  • Pew Survey: Language models tend to reflect the opinions of California, Stanford, and Palo Alto residents.

13. Other Applications:

  • Exoskeletons: Using preference queries to calibrate exoskeletons.
  • Metric Elicitation: Getting pairwise preferences from humans to tune the preference of the learning model.
  • Inverse Decision Theory: Figuring out how to think of trade-offs in classification problems.
  • Recommendation Systems: Modeling and capturing human preferences over items.
  • Reinforcement Learning: Using RL algorithms from human demonstrations.

14. Rationality Assumption:

  • Rational Human: Most of the work will assume some kind of rational human with a deterministic reward function.

15. Conclusion:

  • The course aims to provide a broad overview of machine learning from human preferences, covering various applications, algorithms, and ethical considerations. It emphasizes the importance of understanding and modeling human preferences to improve AI systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.