CIC Statically Generated Graphs for Malware Analysis (CIC-SGG-2024) by Griffin Higgins

THE SUMMARYAI-generated

Key Concepts

  • Malware Analysis: Identifying and understanding malicious software.
  • Static Analysis: Analyzing code without executing it.
  • Graphs/Networks: Data structures with nodes (vertices) and edges representing relationships.
  • Program Graphs: Representations of programs as graphs (Control Flow Graphs, Function Call Graphs).
  • Graph Learning: Applying machine learning, particularly deep learning (Graph Neural Networks), to graphs.
  • Graph Embeddings: Numerical representations of graphs or nodes within graphs.
  • Explainable Malware Detection: Identifying which parts of a program contribute to a malicious classification.
  • Control Flow Graph (CFG): A program graph where nodes represent sequences of non-branching assembly instructions, and edges represent control flow transitions.
  • Function Call Graph (FCG): A program graph where nodes represent functions, and edges represent function calls.

CIC Statically Generated Graphs for Malware Analysis

Introduction

Griffin Higgins, a PhD student and cyber security software developer at the Canadian Institute for Cyber Security (CIC), presents the CIC statically generated graphs data set for malware analysis. The goal is to provide researchers with pre-processed graph representations of malware samples to facilitate research in graph learning and malware detection.

Background

To understand the data set, several key concepts are explained:

  • Malware: Programs designed to cause harm or violate confidentiality, integrity, and availability (CIA). Examples include ransomware and worms.
  • Malware Detection: The task of classifying a program as malicious or benign, ideally with explanations for the classification.
  • Static Analysis: Analyzing a program's code without executing it to determine its behavior.
  • Graphs: Data structures consisting of nodes (vertices) and edges. Edges can be directed or undirected. Graphs can have attributes associated with nodes and edges.
  • Program Graphs: Representing programs as graphs. Two common types are:
    • Control Flow Graphs (CFGs): Nodes represent sequences of non-branching assembly instructions (basic blocks), and edges represent control flow transitions.
    • Function Call Graphs (FCGs): Nodes represent functions, and edges represent function calls.
  • Graph Learning: Applying deep learning techniques, specifically Graph Neural Networks (GNNs), to graphs to learn representations and make predictions. The goal is to move beyond signature-based detection and understand the behavior of malicious programs.
    • Message Passing: Nodes share information with their neighbors iteratively to learn node embeddings that capture the structure and features of the graph.
    • Graph Embeddings: Numerical representations of graphs or nodes within graphs, used for downstream tasks like classification.
  • Graph Learning Tasks:
    • Graph-level tasks: Classifying the entire graph (e.g., malicious or benign).
    • Node-level tasks: Classifying individual nodes within the graph.
    • Edge-level tasks: Predicting the existence or properties of edges.

The CIC Data Set

Motivation

The data set was created because the presenter's previous work on explainable malware detection using graph learning encountered challenges due to the large size of the generated graphs. The data set aims to provide researchers with pre-processed graphs, saving them the effort of setting up a complex pipeline.

Creation Process

The data set creation process involves several steps:

  1. Malicious Binaries: Obtaining malicious and benign PE (Portable Executable) files from various data sets, including:
    • Dyke data set (open)
    • PE for machine learning malware detection data set (open)
    • Bodmos data set (access requires request)
    • Focus is on x86 architecture.
  2. Graph Generation: Using tools like angr to generate CFGs and FCGs from the binary samples. These graphs are represented as NetworkX objects.
    • These graphs contain raw features and are useful for malware analysts.
    • These graphs are not used for training or testing directly; they are used to create embedding graphs.
  3. Embedding Graph Creation: Generating embedding graphs from the attribute graphs.
    • Node features are extracted and used to create node embeddings (numerical representations).
    • These embeddings are used for message passing in GNNs.
    • PyTorch Geometric is used to create the graphs for training and testing.
    • These graphs contain node embeddings and positions, allowing for pruning and referencing back to the original graph.
  4. Graph Pruning (Optional): Applying graph pruning algorithms to reduce the size of the graphs. These algorithms are not included in the data set but can be applied by users.
  5. Explanation Generation: Identifying the nodes and edges that contribute most to a given prediction.

Data Set Contents

The CIC data set includes:

  • Attribute Graphs: CFGs and FCGs with raw features.
  • Embedding Graphs: Graphs with node embeddings, used for training and testing GNNs.
  • Explained Graphs: Graphs with information about which nodes and edges contributed to a prediction.

All three types of graphs have the same structure for a given sample, but the features within each type differ.

Graph Visualization

The presentation includes visualizations of FCGs and CFGs. The CFGs, in particular, can be very large, with some samples reaching up to 300,000 nodes. This highlights the importance of graph reduction techniques.

Data Set Location

The data set can be downloaded from the Canadian Institute for Cyber Security website under the "Data Sets" tab. The data is organized into directories for attribute graphs, embedding graphs, and explained graphs. Individual samples can also be downloaded separately.

Q&A Highlights

  • Largest Challenge: The sheer size of the graphs and the computational complexity of graph algorithms. Igraph library in Python is recommended for faster graph processing.
  • Research Opportunities:
    • Analyzing the explanation data.
    • Improving node embeddings.
    • Exploring different graph learning models.
    • Considering edge-level tasks.
    • Investigating temporal graph learning.
  • Industrial Applications: Malware detection, addressing the limitations of signature-based detection by focusing on program behavior.

Conclusion

The CIC statically generated graphs data set provides a valuable resource for researchers in malware analysis and graph learning. It offers pre-processed graph representations of malware samples, enabling the development of more effective and explainable malware detection techniques. The data set includes attribute graphs, embedding graphs, and explained graphs, providing a comprehensive set of resources for various research directions. The presenter emphasizes the importance of graph reduction techniques due to the large size of the generated graphs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.