Gaussian Naive Bayes From Scratch in Python (Mathematical)

NeuralNineAbout 5 min readJul 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gaussian Naive Bayes: A classification algorithm based on Bayes' theorem, assuming features are normally distributed (Gaussian) and independent (naive).
  • Bayes' Theorem: A mathematical formula for calculating conditional probabilities.
  • Prior Probability: The probability of a class occurring without any prior knowledge of the data.
  • Likelihood: The probability of observing the data given a specific class.
  • Naive Assumption: The assumption that features are independent of each other, given the class.
  • Gaussian Distribution (Normal Distribution): A probability distribution characterized by a bell-shaped curve, defined by its mean and variance.
  • Mean: The average value of a set of numbers.
  • Variance: A measure of how spread out a set of numbers is.
  • Underflow: A condition where a numerical value becomes too small to be represented accurately by a computer.
  • Logarithm: A mathematical function that reverses exponentiation, used here to avoid underflow by converting multiplication to addition.
  • Argmax: A function that returns the argument (class) that maximizes a given expression (probability).

Theory and Mathematics of Gaussian Naive Bayes

The video explains the theory and mathematics behind the Gaussian Naive Bayes classifier. The core idea is to use Bayes' theorem to calculate the probability of a data point belonging to a particular class, given its features.

Bayes' Theorem:

  • P(Y=k | X) = [P(Y=k) * P(X | Y=k)] / P(X)
    • P(Y=k | X): Posterior probability of class k given data X.
    • P(Y=k): Prior probability of class k.
    • P(X | Y=k): Likelihood of observing data X given class k.
    • P(X): Probability of observing data X (ignored for comparison purposes).

Naive Assumption:

  • The algorithm makes a "naive" assumption that the features are independent of each other, given the class. This simplifies the calculation of the likelihood:
    • P(X | Y=k) = Product of P(Xj | Y=k) for all features j.

Gaussian Distribution:

  • The algorithm assumes that each feature follows a Gaussian (normal) distribution within each class.
  • The probability of a feature value Xj given class k is calculated using the Gaussian probability density function:
    • P(Xj | Y=k) = (1 / sqrt(2*pi*variance_kj)) * exp(-(Xj - mean_kj)^2 / (2*variance_kj))
    • mean_kj: Mean of feature j for class k.
    • variance_kj: Variance of feature j for class k.

Underflow Prevention:

  • Multiplying many small probabilities can lead to underflow, causing numerical instability. To avoid this, the algorithm uses logarithms:
    • log(P(Y=k | X)) is proportional to log(P(Y=k)) + Sum of log(P(Xj | Y=k))
  • The logarithm of the Gaussian probability density function is:
    • -0.5 * log(2*pi*variance_kj) - (Xj - mean_kj)^2 / (2*variance_kj)

Prediction:

  • The algorithm calculates the posterior probability for each class and predicts the class with the highest probability (argmax).

Implementation in Python with NumPy

The video demonstrates how to implement the Gaussian Naive Bayes classifier from scratch using Python and NumPy.

Class Structure:

  • GaussianNB class:
    • fit(X, Y) method: Trains the classifier on the input data X and labels Y.
    • predict(X) method: Predicts the class labels for new data X.
    • _log_gaussian(X) method: Calculates the log probabilities based on the Gaussian distribution.

fit(X, Y) Method:

  1. Convert data to NumPy arrays: X = np.asarray(X), Y = np.asarray(Y).
  2. Determine unique classes: self.classes_ = np.unique(Y).
  3. Get the number of classes and features: n_classes = len(self.classes_), n_features = X.shape[1].
  4. Initialize arrays to store means, variances, and priors:
    • self.means_ = np.zeros((n_classes, n_features))
    • self.variances_ = np.zeros((n_classes, n_features))
    • self.priors_ = np.zeros(n_classes)
  5. Iterate through each class:
    • Extract data for the current class: X_k = X[Y == k].
    • Calculate the mean for each feature: self.means_[index] = np.mean(X_k, axis=0).
    • Calculate the variance for each feature: self.variances_[index] = np.var(X_k, axis=0).
    • Calculate the prior probability: self.priors_[index] = X_k.shape[0] / X.shape[0].

_log_gaussian(X) Method:

  1. Calculate the first part of the Gaussian log probability: num = -0.5 * ((X - self.means_) ** 2) / self.variances_.
  2. Calculate the second part of the Gaussian log probability: log_prob = num - 0.5 * np.log(2 * np.pi * self.variances_).
  3. Sum the log probabilities across features: return log_prob.sum(axis=1).

predict(X) Method:

  1. Convert data to a NumPy array: X = np.asarray(X).
  2. Calculate the log likelihoods using the _log_gaussian method: log_likelihood = self._log_gaussian(X).
  3. Calculate the log prior probabilities: log_prior = np.log(self.priors_).
  4. Find the class with the maximum log probability: return self.classes_[np.argmax(log_likelihood + log_prior, axis=1)].

Evaluation and Results

The video uses the breast cancer dataset from scikit-learn to evaluate the implemented classifier.

Data Loading and Preprocessing:

  • load_breast_cancer(return_X_y=True): Loads the breast cancer dataset, returning the data (X) and labels (Y).
  • train_test_split(X, Y, test_size=0.2): Splits the data into training and testing sets with an 80/20 ratio.

Evaluation Metrics:

  • accuracy_score(Y_test, Y_pred): Calculates the accuracy of the classifier by comparing the predicted labels (Y_pred) to the true labels (Y_test).

Results:

  • The implemented Gaussian Naive Bayes classifier achieves an accuracy of around 92-96% on the breast cancer dataset.

Conclusion

The video provides a comprehensive explanation of the Gaussian Naive Bayes classifier, including its theoretical foundations, mathematical formulas, and a practical implementation in Python using NumPy. The implementation is evaluated on the breast cancer dataset, demonstrating its effectiveness as a classification algorithm. The use of logarithms to prevent underflow is a key aspect of the implementation. The video emphasizes the importance of understanding the underlying mathematics before implementing machine learning algorithms from scratch.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.