System Modeling and Property Specification: A Deep Dive
Key Concepts:
- System Modeling: Representing a system's behavior through mathematical models, including its agent, sensor, and environment.
- Observation Model: The probability of observing a particular output given a specific state.
- Model Class: A family of mathematical functions used to represent the system's behavior (e.g., linear, Gaussian).
- Parameter Learning: Estimating the parameters of a model class from observed data.
- Maximum Likelihood Estimation (MLE): Finding the parameter values that maximize the likelihood of observing the given data.
- Bayesian Parameter Learning: Maintaining a probability distribution over all possible parameter values, incorporating prior beliefs and observed data.
- Prior Distribution: A probability distribution representing our belief about the parameter values before observing any data.
- Likelihood Model: The probability of observing the data given specific parameter values.
- Posterior Distribution: A probability distribution representing our updated belief about the parameter values after observing the data.
- Probabilistic Programming: A programming paradigm that allows us to draw samples from the posterior distribution even when it's difficult to compute analytically.
- Markov Chain Monte Carlo (MCMC): A class of algorithms used in probabilistic programming to sample from probability distributions.
- Conjugate Prior: A prior distribution that, when combined with a specific likelihood model, results in a posterior distribution within the same family.
- Validation Set: A separate set of data used to assess the performance of a trained model and prevent overfitting.
- Cross-Validation: A technique for evaluating a model's performance by training and testing it on different subsets of the data.
- KL Divergence: A measure of how different two probability distributions are.
- KS Statistic: A measure of the maximum distance between two cumulative distribution functions (CDFs).
- QQ Plot (Quantile-Quantile Plot): A plot comparing the quantiles of two distributions.
- Calibration Plot: A plot comparing the predicted probabilities of a model to the observed frequencies.
- Property Specification: Formally defining what a system is supposed to do.
- Metric: A function that maps system behavior to a real number.
- Specification: A function that maps system behavior to a Boolean value (true or false).
- Risk Metric: A metric for which higher values indicate worse outcomes.
- Value at Risk (VAR): The highest value that a risk metric is guaranteed not to exceed with a given probability (alpha).
- Conditional Value at Risk (CVAR): The average of all risk values above the VAR.
- Pareto Optimality: A state where it's impossible to improve one metric without making another metric worse.
- Pareto Frontier: The set of all Pareto optimal designs.
- Composite Metric: A single metric that combines multiple individual metrics.
- Weighted Sum: A composite metric calculated by assigning weights to individual metrics and summing them.
- Goal Distance Metric: A composite metric calculated by measuring the distance from each design to a desired goal point (utopia point).
I. System Modeling: Building Representations of Reality
A. Concrete Example: Aircraft Altitude Sensor
- Scenario: Modeling an aircraft's altitude sensor.
- State: The aircraft's actual height above the ground (h).
- Observation: The sensor's measured height (h-hat), which is a noisy version of the actual height.
- Data Collection: Pairs of (h, h-hat) values are collected.
- Observation Model: The probability of observing h-hat given the actual height h: P(h-hat | h).
- Model Class Selection:
- The relationship between h and h-hat appears linear with added noise.
- A conditional Gaussian model is chosen: h-hat ~ f_theta(h) + noise.
- A linear function is assumed: f_theta(h) = theta_1 * h + theta_2.
- Parameter Optimization:
- Maximum likelihood estimation (MLE) is used to find the optimal parameters (theta_1, theta_2, sigma^2).
- Example: theta_1 = 1, theta_2 = 0, sigma^2 = 0.05.
- Non-Linear Example: If the relationship between h and h-hat is non-linear, a more expressive model class is needed.
B. Bayesian Parameter Learning: Embracing Uncertainty
- Motivation: Overcoming indecisiveness by maintaining a distribution over all possible parameter values.
- Goal: Modeling P(theta | D), the probability of the parameters given the data.
- Bayes' Rule: P(theta | D) = [P(D | theta) * P(theta)] / P(D)
- Likelihood Model: P(D | theta), the probability of the data given the parameters (same as in MLE).
- Prior: P(theta), our belief about the parameter values before seeing any data.
- Posterior: P(theta | D), our updated belief about the parameter values after seeing the data.
- Computational Challenge: Computing the posterior analytically is often difficult due to the integral in the denominator.
- Probabilistic Programming Solution:
- We can compute the numerator of Bayes' rule: P(D | theta) * P(theta).
- Probabilistic programming tools (e.g., Turing.jl in Julia) can draw samples from the posterior distribution based on evaluations of the numerator.
- Process:
- Define the likelihood model.
- Define the prior distribution.
- Choose a sampler (e.g., No-U-Turn Sampler - NUTS).
- Specify the number of samples to draw.
- Use a probabilistic programming framework to draw samples from the posterior.
- Example: Learning parameters (theta_1, theta_2, theta_3) for a model.
- Prior: Multivariate normal distribution centered at 0 with a large variance.
- Sampler: NUTS.
- Number of samples: 1000.
- Increasing the number of data points leads to narrower posterior distributions, indicating greater confidence in the parameter estimates.
C. Conjugate Priors: Analytical Solutions
- Frisbee Flipping Experiment:
- Goal: Estimate theta, the probability that two flipped frisbees land facing the same direction.
- Data: A sequence of "same" or "different" outcomes from multiple frisbee flips.
- n: Number of times the frisbees landed facing the same way.
- m: Total number of frisbee flips.
- Likelihood Model: Binomial distribution: P(D | theta) = theta^n * (1 - theta)^(m - n).
- Prior Distribution: Beta distribution with parameters alpha and beta.
- Posterior Distribution: Beta distribution with updated parameters: Beta(alpha + n, beta + m - n).
- Conjugate Prior Definition: A prior distribution is a conjugate prior for a likelihood model if the posterior distribution is in the same family as the prior.
- The beta distribution is the conjugate prior for the Bernoulli distribution.
- Intuition: The alpha and beta parameters of the beta distribution represent prior beliefs about the probability of the frisbees landing the same way.
- Example: Starting with a uniform prior (Beta(1, 1)) and updating it based on observed data from frisbee flips.
D. Parameter Learning: Generalization and Validation
- Overfitting: A model that fits the training data too well may not generalize well to new data.
- Validation Set: A separate set of data used to assess the model's performance and prevent overfitting.
- Cross-Validation: A technique for evaluating a model's performance by training and testing it on different subsets of the data.
- Model Validation: Comparing features of the model to features of the data to ensure they match.
- Visual Diagnostics:
- Probability Density Comparison: Comparing the probability density of the data to the probability density of the model.
- Cumulative Density Comparison: Comparing the cumulative density functions (CDFs) of the data and the model.
- QQ Plot (Quantile-Quantile Plot): Plotting the quantiles of the data against the quantiles of the model.
- Calibration Plot: Plotting the predicted probabilities of the model against the observed frequencies.
- Summary Metrics:
- KL Divergence: A measure of how different two probability distributions are.
- KS Statistic: A measure of the maximum distance between two CDFs.
- Maximum Calibration Error: The maximum distance between the calibration plot and the line y = x.
- Expected Calibration Error: The average distance between the calibration plot and the line y = x.
- Comparing Multiple Features:
- Hand-designing a single feature that combines multiple features.
- Extending summary metrics (e.g., KL divergence) to multiple dimensions.
- Conditional Distributions: Partitioning the conditioning variable into bins and making comparisons in each bin individually.
- Expert Knowledge: Using expert knowledge to validate models, such as a Turing-like test where experts try to distinguish between model outputs and real data.
II. Property Specification: Defining System Requirements
A. Metrics vs. Specifications:
- Property Specification: Formally specifying what a system is supposed to do.
- Metric: A function that maps system behavior to a real number.
- Specification: A function that maps system behavior to a Boolean value (true or false).
- Relationship: Specifications can be derived from metrics, and metrics can be derived from specifications.
- Example: Aircraft Collision Avoidance System:
- Metric: Miss distance between aircraft.
- Specification: Miss distance must be greater than 50 meters.
- Derived Metric: Probability that the miss distance is greater than 50 meters.
B. Metrics for Stochastic Systems:
- Focus: Metrics that summarize distributions of system behavior.
- Process:
- Apply a metric to each trajectory in the distribution.
- Obtain a distribution over the metric.
- Summarize the distribution over the metric.
- Summarization Techniques:
- Expected Value: The mean or average of the distribution.
- Variance: A measure of the spread of the distribution.
- Risk Metrics: Metrics for which higher values indicate worse outcomes.
- Value at Risk (VAR): The highest value that the risk is guaranteed not to exceed with probability alpha.
- Conditional Value at Risk (CVAR): The average of all risk values above the VAR.
- Intuition: VAR and CVAR provide conservative measures of risk, focusing on worst-case scenarios.
C. Composite Metrics: Balancing Trade-offs
- Motivation: Combining multiple metrics that may be at odds with each other.
- Example: Collision Avoidance System:
- Alert Rate: The frequency of alerts issued by the system.
- Collision Rate: The frequency of collisions.
- Pareto Optimality: A system design is Pareto optimal if it's impossible to improve one metric without making another metric worse.
- Pareto Frontier: The set of all Pareto optimal designs.
- Composite Metric Types:
- Weighted Sum: Assigning weights to individual metrics and summing them.
- Goal Distance Metric: Measuring the distance from each design to a desired goal point (utopia point).
- Weighted Exponential Sum Metric: Combining weights and a goal distance metric.
Conclusion:
This lecture provides a comprehensive overview of system modeling and property specification. It covers techniques for building mathematical models of systems, estimating model parameters, validating model accuracy, and formally defining system requirements. The lecture emphasizes the importance of considering uncertainty, balancing trade-offs, and focusing on worst-case scenarios in safety-critical applications. The concepts and techniques presented are essential for designing and validating reliable and robust systems.
AI summaries can miss context or contain errors. Check important details against the original video.