遇见数据集

FCG-MFD: Benchmark Function Call Graph-Based Dataset for Malware Family Detection

收藏
Figshare2024-10-29 更新2026-04-28 收录
官方服务:

资源简介:

Cyber crimes related to malware families are on the rise. This growth persists despite the prevalence of various antivirus software and approaches for malware detection and classification. Security experts have implemented Machine Learning (ML) techniques to identify these cyber-crimes. However, these approaches demand updated malware datasets for continuous improvements amid the evolving sophistication of malware strains. Thus, we present the FCG-MFD, a benchmark dataset with extensive Function Call Graphs (FCG) for malware family detection. This dataset guarantees resistance against emerging malware families by enabling security systems. Our dataset has two sub-datasets (FCG & Metadata) (1,00,000 samples) from VirusSamples, Virusshare, VirusSign, theZoo, Vx-underground, and MalwareBazaar curated using FCGs and metadata to optimize the efficacy of ML algorithms. We suggest a new malware analysis technique using FCGs and graph embedding networks, offering a solution to the complexity of feature engineering in ML-based malware analysis. Our approach to extracting semantic features via the Natural Language Processing (NLP) method is inspired by tasks involving sentences and words, respectively, for functions and instructions. We leverage a node2vec mechanism-based graph embedding network to generate malware embedding vectors. These vectors enable automated and efficient malware analysis by combining structural and semantic features. We use two datasets (FCG & Metadata) to assess FCG-MFD performance. F1-Scores of 99.14% and 99.28% are competitive with State-of-the-art (SOTA) methods.

与恶意软件家族相关的网络犯罪呈持续上升态势。尽管当前已有各类杀毒软件及恶意软件检测与分类方法,但此类犯罪的增长势头并未得到有效遏制。安全领域专家已引入机器学习(Machine Learning, ML)技术以识别此类网络犯罪。然而,随着恶意软件变体的复杂度不断提升,上述方法需要依托最新的恶意软件数据集来实现持续优化。为此,我们提出FCG-MFD数据集——一款包含海量函数调用图(Function Call Graphs, FCG)的恶意软件家族检测基准数据集。该数据集可助力安全系统抵御新兴恶意软件家族。 本数据集包含两个子数据集(函数调用图与元数据),共计100000条样本,数据来源于VirusSamples、Virusshare、VirusSign、theZoo、Vx-underground及MalwareBazaar,并通过函数调用图与元数据进行精选整理,以优化机器学习算法的应用效能。 我们提出了一种基于函数调用图与图嵌入网络的新型恶意软件分析技术,旨在解决基于机器学习的恶意软件分析中特征工程的复杂性难题。我们提出的基于自然语言处理(Natural Language Processing, NLP)的语义特征提取方法,分别借鉴了针对句子与词汇的处理任务思路,以适配函数与指令的分析需求。我们采用基于node2vec机制的图嵌入网络生成恶意软件嵌入向量,该向量可结合结构特征与语义特征,实现自动化且高效的恶意软件分析。 我们通过两个子数据集(函数调用图与元数据)对FCG-MFD的性能进行评估,其F1值分别达到99.14%与99.28%,性能可与当前最优(State-of-the-art, SOTA)方法相媲美。

创建时间:
2024-10-29
二维码
社区交流群
二维码
科研交流群
商业服务