Stanford CS221 | Autumn 2025 | Lecture 12: Bayesian Networks I

By Stanford Online

Share:

Key Concepts

  • Bayesian Networks: Graphical models representing a joint probability distribution over a set of random variables using a directed acyclic graph (DAG).
  • Joint Distribution: A probability table assigning a probability to every possible assignment of values to all random variables in a system.
  • Marginalization: The process of "collapsing" or summing out variables from a joint distribution to focus on a subset of variables.
  • Conditioning: Updating probabilities based on observed evidence (e.g., $P(A|B)$), involving selection of relevant rows and renormalization.
  • Explaining Away: A reasoning pattern where observing an effect and one cause reduces the probability of an alternative cause.
  • Probabilistic Programming: Defining joint distributions through code (programs) that generate samples, allowing for inference via simulation.
  • Rejection Sampling: An approximate inference method that generates samples from a model and keeps only those that match the observed evidence.
  • Einsum (Einstein Summation): A notation/framework used to perform tensor operations (marginalization, multiplication) efficiently on probability tables.

1. Foundations of Bayesian Networks

The lecture transitions from reactive machine learning (predicting actions from states) to model-based reasoning, where an agent understands the underlying mechanics of the world. Bayesian networks serve as the representation for this world model.

  • The Four-Step Construction:
    1. Define Variables: Identify the attributes of the world (e.g., Burglary, Earthquake, Alarm).
    2. Define Edges: Create a directed graph representing dependencies (e.g., Burglary $\rightarrow$ Alarm).
    3. Local Conditional Distributions: Define the probability of each node given its parents (e.g., $P(A|B, E)$).
    4. Joint Distribution: Multiply all local conditional distributions together to form the complete joint distribution.

2. Probabilistic Inference

Once the joint distribution is established, it acts as a "database" for answering queries.

  • Methodology:
    • Selection: Filter the joint distribution to include only rows consistent with the evidence.
    • Marginalization: Sum out variables not involved in the query or evidence.
    • Renormalization: Divide the selected probabilities by the total probability of the evidence to ensure the resulting distribution sums to 1.
  • Technical Note: Probability tables are treated as tensors. Operations like marginalization and conditioning are performed using einsum to handle high-dimensional data efficiently.

3. The "Explaining Away" Phenomenon

A key argument presented is that Bayesian networks naturally capture complex reasoning patterns.

  • Case Study (Alarm/Earthquake/Burglary): If an alarm sounds, the probability of a burglary increases. However, if you subsequently learn there was an earthquake, the probability of a burglary decreases.
  • Logic: The earthquake "explains away" the alarm, reducing the need to attribute the alarm to a burglary. This occurs even if the two causes (earthquake and burglary) are independent.

4. Probabilistic Programming and Sampling

The instructor introduces an alternative, code-based representation of Bayesian networks.

  • Probabilistic Programs: Instead of static tables, models are defined as functions that return random samples.
  • Rejection Sampling:
    • Process: Generate many samples from the program; discard those that do not match the evidence; compute a histogram of the remaining samples.
    • Pros/Cons: It is conceptually simple and works for any model, but it is computationally inefficient for rare events (where most samples are rejected).
    • Convergence: As the number of samples approaches infinity, the estimate converges to the true probability.

5. Comparison: Traditional ML vs. Bayesian Networks

| Feature | Traditional ML (Classifiers) | Bayesian Networks | | :--- | :--- | :--- | | Direction | Input $\rightarrow$ Output | Output $\rightarrow$ Input (Generative) | | Missing Data | Often requires imputation | Handled naturally via marginalization | | Prior Knowledge | Hard to incorporate | Easily encoded in the graph structure | | Interpretability | Often "black box" | High (intermediate variables are meaningful) |

6. Notable Quotes

  • "The joint distribution is like a SQL database, and probabilistic inference is essentially doing SQL queries on that database."
  • "If you observe variables that are downstream of things that are independent, that makes the variables upstream not independent anymore."

Synthesis

Bayesian networks provide a rigorous mathematical framework for reasoning under uncertainty. By defining local dependencies, one can construct a global joint distribution that allows for flexible inference—answering any query given any evidence. While exact inference can be computationally expensive (NP-hard), techniques like probabilistic programming and rejection sampling offer intuitive, albeit approximate, ways to reason about complex systems. The shift from discriminative modeling to generative Bayesian modeling allows for better handling of missing data, incorporation of prior knowledge, and causal reasoning.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video