Stanford AA228V I Validation of Safety Critical Systems I Falsification through Optimization

Unknown AuthorAbout 5 min readApr 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Falsification: Systematically finding failures of a system.
  • Direct Falsification (Direct Sampling): Directly sampling the system to find failures.
  • Rare Failure Events: Failures that occur with very low probability.
  • Probability of Failure (P fail): The probability that a system will fail.
  • Geometric Distribution: A probability distribution that describes the probability of the number of trials needed for one success.
  • Disturbances: A way to formulate the problem that allows us to systematically search for failures by taking control of the sources of randomness in the system.
  • Disturbance Distribution: The distribution from which disturbances are sampled.
  • Nominal Trajectory Distribution: Represents the distribution over trajectories that we would expect to observe when we deploy the system in the real world.
  • Fuzzing: Inputting a different distribution into the rollout, not the nominal distribution, but one that we think is more likely to actually cause failures or result in rollouts that are failures.
  • Optimization: Systematically searching over initial states and disturbances or sequence of disturbances to find failure trajectories.
  • Robustness: A metric that represents closeness to failure.

Falsification and its Importance

  • Falsification is the process of systematically finding failures of a system.
  • The primary goal of falsification is to inform future design decisions by uncovering potential issues before real-world deployment.
  • Failures identified through falsification can lead to:
    • Enhancements in system sensors.
    • Updates to the agent's policy.
    • Revisions of system requirements.
    • Adaptation of human operator training.
    • Recognition of system limitations.
    • Abandonment of the project if necessary.

Direct Falsification (Direct Sampling)

  • Direct falsification is the simplest approach, involving direct sampling of the system to find failures.
  • The algorithm takes a depth (trajectory length) and a number of samples as input.
  • It performs multiple rollouts (simulations) and filters out the trajectories that result in failures.
  • Direct falsification may perform poorly for systems with rare failure events.
  • The probability of sampling a failure on the k-th rollout follows a geometric distribution.
  • The expected number of simulations required to find a single failure event is 1/P fail, where P fail is the probability of failure.
  • For systems with very low P fail (e.g., 10^-9 for aircraft collision avoidance), direct falsification may require an impractical number of simulations.

Disturbances: A Powerful Formulation

  • Disturbances are introduced as a way to systematically search for failures by taking control of the sources of randomness in the system.
  • To incorporate disturbances, all three components of the system (agent, environment, and sensor model) are rewritten.
  • For example, the observation model is rewritten as O(s, Xo), where s is the state and Xo is a disturbance sampled from a disturbance distribution.
  • This separates the deterministic and stochastic components of the system.
  • The disturbance distribution is typically simple and well-defined, allowing for the computation of likelihoods.
  • The agent, environment, and sensor model are all rewritten in terms of disturbances and deterministic functions.
  • All three components are packaged into a single disturbance (x) and a single disturbance distribution.
  • The step function is rewritten to incorporate disturbances, allowing for explicit control over the randomness in the system.

Trajectory Distributions

  • A trajectory distribution defines the distribution over possible trajectories for a system.
  • It consists of a distribution over initial states and a disturbance distribution that takes in a time.
  • The depth of the trajectory is also specified.
  • Sampling from a trajectory distribution involves sampling an initial state and then repeatedly taking a step using the disturbance distribution.
  • The nominal trajectory distribution represents the distribution over trajectories that we would expect to observe when we deploy the system in the real world.
  • Nominal models are typically derived from data and expert knowledge of the real world.

Fuzzing: Introducing a Different Distribution

  • Fuzzing involves inputting a different distribution into the rollout, not the nominal distribution, but one that we think is more likely to actually cause failures or result in rollouts that are failures.
  • This typically requires some domain knowledge to figure out.
  • For example, in the inverted pendulum, increasing the amount of perception noise can lead to more failures.
  • It is important to measure likelihood from the nominal distribution, even when sampling from the fuzzing distribution.
  • The goal is to find failures that are the result of likely events, not just any failure.

Optimization: Systematically Searching for Failures

  • Optimization involves systematically searching over initial states and disturbances to find failure trajectories.
  • The decision variables are the initial state (S) and the sequence of disturbances (boldface X).
  • The objective function represents closeness to failure and is minimized over the decision variables.
  • The rollout function is used as a constraint, ensuring that the trajectory comes from a rollout with the specified initial state and disturbances.
  • The objective function can be set to the robustness metric, which represents closeness to failure.
  • A potential problem with optimization is that it may find very unlikely trajectories.
  • To address this, likelihood should be incorporated into the objective function.

Conclusion

The lecture provides a comprehensive overview of falsification algorithms, starting with basic direct sampling and progressing to more sophisticated techniques like fuzzing and optimization. The concept of disturbances is introduced as a powerful tool for formulating the problem and enabling systematic search for failures. The importance of considering likelihood when searching for failures is emphasized, particularly in the context of optimization. The lecture also provides practical guidance for project one, recommending fuzzing as a starting point and optimization for improving leaderboard scores.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.