Leveraging AI Knowledge Distillation for Deployable Cybersecurity Defense Systems by Mahdi Rabbani

By Canadian Institute for Cybersecurity (CIC)

Share:

Leveraging AI Knowledge Distillation for Deployable Cyber Security Defense Systems: A Detailed Summary

Key Concepts:

  • Knowledge Distillation: A model compression technique where a smaller “student” model learns from the “soft probabilities” (knowledge) of a larger, pre-trained “teacher” model.
  • Dark Knowledge: The information contained within the probability distribution of a model’s predictions, beyond the single predicted class. Reveals relationships between classes.
  • Temperature (T): A parameter used in the softmax function to control the smoothness of the probability distribution. Higher temperatures reveal more dark knowledge.
  • KL Divergence: A measure of how one probability distribution differs from a second, reference probability distribution. Used to quantify the difference between student and teacher outputs.
  • Ensemble Learning: Combining multiple machine learning models to improve overall performance. (Bagging, Boosting, Stacking)
  • Softmax: A function that converts a vector of numbers into a probability distribution.
  • Argmax: A function that returns the index of the maximum value in an array.
  • MobileBERT & BiLSTM: Specific model architectures used in the presented case study – MobileBERT as the teacher and BiLSTM as the student.

1. Introduction & Background

Dr. Madi Rabani Aron from the Canadian Institute for Cyber Security presented a webinar on leveraging AI knowledge distillation for deployable cyber security defense systems. The motivation stems from the need to translate AI research into practical, real-world applications, particularly in areas like malware detection, analysis, and attack classification. Dr. Rabani Aron has been working with AI in cybersecurity for 10 years. The presentation aimed to bridge the gap between theoretical AI concepts and their deployment in industry settings.

2. Traditional Machine Learning & its Limitations

The presentation began with a review of traditional machine learning workflows. A dataset is divided into training and testing sets, with labels assigned to training samples. The model learns to predict these labels. However, the standard approach focuses solely on the “winning class” (highest probability) during prediction, discarding valuable information about the model’s confidence in other classes. This loss of information is a key limitation. This applies to various models – KNN, Neural Networks, SVMs, Decision Trees, Random Forests – all ultimately reducing to a single class prediction.

3. Ensemble Learning & the Value of Soft Averaging

To address the limitations of single models, ensemble learning was introduced. This involves training multiple diverse models and combining their predictions. The goal is to minimize correlation between model errors. Common ensemble methods include Bagging (Random Forest), Boosting (XGBoost), and Stacking. While averaging predictions improves performance, it still typically focuses on the winning class, losing the nuanced information contained in the full probability distribution. The concept of “soft averaging” was highlighted – utilizing the full probability distribution before applying the argmax function. This preserves more knowledge about relationships between classes.

4. Introducing Knowledge Distillation & Dark Knowledge

The core of the presentation focused on knowledge distillation. The key idea is that the probability distribution before the argmax function (the “soft probabilities”) contains valuable “dark knowledge” about how the model relates different classes. For example, a model might recognize a visual similarity between a dog and a cat, even if it ultimately classifies the image as a dog. This relationship is lost when only the winning class is considered.

An example was given: a model correctly identifying a "dog" but also showing a relatively high probability for "cat," indicating a perceived similarity. This is dark knowledge. This is particularly useful in image generation tasks, where understanding relationships between concepts is crucial.

5. Teacher-Student Framework & Mathematical Formulation

Knowledge distillation involves a “teacher” model (typically a large, complex model) and a “student” model (typically smaller and more efficient). The student learns from both the ground truth labels and the soft probabilities generated by the teacher.

  • Temperature (T): A temperature parameter is applied to the softmax function to control the smoothness of the probability distribution. Higher temperatures (3-5 are recommended) reveal more dark knowledge by flattening the distribution.
  • KL Divergence: KL divergence is used to measure the difference between the student’s and teacher’s probability distributions, guiding the student’s learning process.
  • Loss Function: The student’s training involves a combined loss function: a distillation loss (based on KL divergence) and a traditional loss (based on ground truth labels). A balancing factor (alpha = 0.5 in the presented case study) controls the relative contribution of each loss component.

6. Case Study: Fishing Email Detection

A practical case study was presented involving the detection of phishing emails.

  • Dataset: A dataset of 1 million samples (500,000 phishing, 500,000 legitimate) was used, augmented with synthetically generated samples using Large Language Models (LLMs).
  • Teacher Model: MobileBERT, a smaller version of BERT, was used as the teacher model.
  • Student Model: A Bidirectional LSTM (BiLSTM) was used as the student model. BiLSTM processes sequential data (text) in both directions (forward and backward) to capture contextual information.
  • Architecture: The student model’s architecture was designed to be similar to the teacher model, facilitating knowledge transfer.
  • Results: The distilled student model (BiLSTM) achieved comparable performance to the teacher model (MobileBERT) while being significantly smaller and faster. The student model had approximately 25 million parameters compared to the teacher’s 110 million. Inference time was reduced from 42 seconds (MobileBERT) to 6 seconds (BiLSTM).

7. Performance Comparison & Analysis

The performance of different models was compared: LSTM, BiLSTM, BiLSTM with attention mechanisms (single-head and multi-head), and the knowledge-distilled BiLSTM. The knowledge-distilled BiLSTM consistently outperformed the other models, demonstrating the effectiveness of the technique. The results showed improvements in accuracy, precision, recall, and F1-score across various training and testing scenarios (original data, generated data, mixed data).

8. Future Directions & Conclusion

Dr. Rabani Aron highlighted the potential of small language models (SLMs) for agentic AI, particularly in resource-constrained environments. Knowledge distillation is crucial for creating these SLMs. Future research directions include:

  • Graph Condensation: Applying knowledge distillation to graph learning for large-scale graphs.
  • Robustness against Adversarial Attacks: Investigating the robustness of distilled models against AI-generated adversarial attacks (e.g., sophisticated phishing emails).
  • Maintaining Teacher Knowledge: Addressing the potential for the student model to “forget” the teacher’s knowledge when trained on new data.

Key Takeaway: Knowledge distillation is a powerful technique for compressing AI models without significant performance loss, enabling deployment in resource-constrained environments like mobile devices and IoT devices. It leverages the “dark knowledge” contained within the soft probabilities of larger models to improve the performance of smaller models. This is particularly relevant for cybersecurity applications where efficient and accurate threat detection is critical.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video