Stanford Seminar - On human-machine interaction games

Unknown AuthorAbout 6 min readMay 12, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Human-Machine Interaction: The study of how humans and machines interact, focusing on predicting and shaping outcomes.
  • Forward and Inverse Models: Internal representations of system dynamics used for control. Feedforward controllers approximate the inverse of forward models.
  • Feedforward and Feedback Control: Control strategies where feedforward anticipates desired actions and feedback corrects errors.
  • Game Theory Equilibria: Nash, Stackelberg, and Conjectural Equilibria, representing different outcomes in human-machine interaction games based on cost functions and learning strategies.
  • Bounded Rationality: The idea that agents (humans and machines) have limitations in their ability to perfectly optimize, leading to local search methods.
  • Cost Functions: Mathematical representations of the goals and priorities of humans and machines, used to model decision-making.
  • Myoelectric Interface: Using muscle activity (EMG) to control a machine, such as a cursor on a screen.
  • Learning Rate (Alpha): A hyperparameter in learning algorithms that determines the step size during adaptation.
  • Reverse Stackelberg Equilibrium (RSE): An outcome where the machine designs an incentive for the human to behave in a way that benefits the machine.
  • Encoder/Decoder: In the context of brain-machine interfaces, the encoder represents the brain's mapping of intention to muscle activity, and the decoder maps muscle activity to machine control.

Models People Learn When Controlling Machines

  • Motivation: The increasing mediation of human interaction with the physical world by machines (teleoperated robots, neural interfaces, assistive devices).
  • Research Goal: To predict and shape the outcomes of human-machine interactions, leading to better-performing, more usable, and preferable devices.
  • Experimental Paradigm: A human controls a remote robot through an interface, tracking a reference signal while disturbances are introduced.
  • System Representation: The system is modeled as a block diagram with interconnected dynamical systems, including human input (u), machine output, reference (r), and disturbances.
  • Human Transformation: The human transformation involves looking at the reference, the error, and producing a control signal (u).
  • Hypotheses: The human transformation includes feedforward and feedback mechanisms.
  • Feedforward Controllers: Approximate the inverse of forward models or machine dynamics.
  • Feedback Mechanisms: Operate on the error between the desired and actual states.
  • Mathematical Representation: The block diagram is transcribed into a system of equations to estimate feedforward (F) and feedback (B) transformations.
  • Experimental Setup: Participants track a reference signal on a screen using a manual or myoelectric interface.
  • Bode Plots: Used to analyze the frequency response of human transformations, revealing surprisingly linear dynamics.
  • Feedforward Approximation: The feedforward transformation approximates the inverse model but with a systematic bias.
  • Adaptation: Humans adapt well to different system dynamics, learning categorically different models for first-order and second-order systems.
  • Myoelectric Interface Comparison: Compared to manual interfaces, myoelectric interfaces show comparable tracking error for first-order systems but better inversion error for second-order systems.
  • Hypothesis for Myoelectric Superiority: The shorter delay in muscle activation compared to manual input may contribute to better feedforward signal implementation.
  • Key Takeaway: Humans learn different models depending on the system they are interacting with, approximately inverting machine dynamics.

Machine Learning from and Interacting with People

  • Flipping the Script: The machine's model (M) becomes a function of the human's actions (H), creating a mathematical game.
  • Mathematical Game: Two decision-making agents (human and machine) with potentially differing or conflicting priorities and imperfect information sharing.
  • Cost Function Minimization: Both agents are hypothesized to make decisions by minimizing their respective cost functions.
  • Bounded Rationality Assumption: Agents do not perfectly globally optimize but use local search methods like gradient descent.
  • Simplest Possible Human-Machine Interaction Game: A scalarized problem where both human and machine have one-dimensional decision variables.
  • Cost Function Prescription: The cost function is prescribed to the user for initial experiments, then left up to the user in later experiments.
  • Multiple Candidate Outcomes: With two cost functions, there are multiple possible outcomes, including the human or machine's global optimum, Nash equilibrium, Stackelberg equilibrium, and conjectural equilibrium.
  • Best Response Functions: Represent the optimal response of one agent to a fixed action of the other agent.
  • Nash Equilibrium: The intersection of best response curves, where neither agent has an incentive to deviate on their own.
  • Stackelberg Equilibrium: Arises when there is an order of play, with one agent (e.g., human) leading and the other (e.g., machine) following.
  • Conjectural Equilibrium: Arises when the machine builds a model of what the human is doing and responds to that, leading to internal models being formed by both players.
  • Experimental Setup: Participants use a mouse or touchscreen to control a cursor horizontally, influencing the height of a vertical bar, while a machine action also influences the bar's height.
  • Crowdsourced Platform (Prolific): Used to collect data quickly and cheaply.
  • Experiment 1: Gradient Descent Adaptation: The machine adapts using gradient descent with varying learning rates (alpha).
  • Results: As the learning rate increases, the system shifts from Nash equilibrium to Stackelberg equilibrium.
  • Takeaway: Results are consistent with the human doing something like gradient descent on their own cost.
  • Experiment 2: Targeting Conjectural Equilibrium: The machine estimates the parameters of a human model and uses it to solve its optimization problem.
  • Conjectural Variations: The machine builds a model of the human, and the human builds a model of the machine, iterating towards a consistent conjectural variations equilibrium.
  • Results: The system shifts from Stackelberg equilibrium to conjectural equilibrium.
  • Takeaway: Strong evidence that the human is learning an internal model for the machine.
  • Experiment 3: Steering the System Outcome: The machine implements a policy gradient algorithm to steer the system towards its global minimum.
  • Results: The machine converges to its global minimum, coercing the human's behavior.
  • Reverse Stackelberg Equilibrium (RSE): The outcome where the machine designs an incentive for the human to behave in a way that benefits the machine.
  • Takeaway: Machine learning algorithms can strongly influence human behavior, even with simple algorithms.
  • Game Theory Equilibria: A wide variety of game theory equilibria can be achieved by the machine choosing a learning algorithm.
  • High-Dimensional Noninvasive Brain-Machine Interface: Using a high-density EMG array to control a cursor on a screen.
  • Cost Function Hypothesis: Both the machine and the brain have a cost function that is a combination of tracking error and effort.
  • Penalty Parameter (Lambda): A hyperparameter that weights the effort term in the cost function.
  • Prediction: The decoder and brain will learn to approximately invert one another.
  • Results: Experimental results are consistent with control theory predictions, including approximate inversion and stability of the closed-loop system.
  • Varying Penalty Parameter: Cranking up the penalty parameter in the machine's cost function reduces decoder effort and increases encoder effort.
  • Takeaway: Machine learning algorithms can not only predict outcomes but also shape them, such as manipulating user effort.

Conclusion

The research demonstrates the power of game theory in understanding and shaping human-machine interactions. By modeling both human and machine behavior as cost function minimization problems, the researchers were able to predict and influence outcomes, including selecting different game theory equilibria and manipulating user effort. The choice of machine learning algorithm plays a crucial role in determining the outcome of these interactions, highlighting the potential for machines to influence human behavior in significant ways. The findings have implications for various applications, including rehabilitation, assistive devices, and brain-machine interfaces.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.