Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches (Replication Package Part 4: Poppler Dataset)
收藏资源简介:
The Replication Package of "Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches" Part 4 (POPPLER Dataset) This repository includes: Code.zip that contains the codes to replicate some parts of this study:a. 1_generate_datasets implements our methodology to generate the datasets.b. 2_run_models runs the ML models during the evaluation.c. 3_result_replication generates charts presented in the paper from the ML evaluation results. Datasets.zip that contain 2 folders:a. original datasets: 1 from NVD Vuldeepecker and 3 extracted from BigVul. b. POPPLER datasets: train, validation, test sets for each time of observation extracted using our methodology from BigVul dataset for project poppler. Pretrained-models.zip that we generated during our evaluation (3 test results for each time point in the timeline [2009-2018]). Results.zip of our evaluation, the folder ALL contains the overall results and other folders are results by model. Documentations INSTALL.pdf : how to install the codes README.pdf: readme file REQUIREMENTS.pdf: hardware and software requirements STATUS.pdf : status for artifact submission LICENSE.pdf: the license of this artifact PAPER.pdf: the camera-ready version of the paper Please refer to the following repositories for the other datasets and pre-trained models: - Part 1 NVD Vuldeeepecker : https://doi.org/10.5281/zenodo.8207883- Part 2 LINUX : https://doi.org/10.5281/zenodo.10960662- Part 3 OPENSSL : https://doi.org/10.5281/zenodo.10966117
《今日之猫,明日之犬:机器学习漏洞检测方法标签时变特性考量》复制包 第四部分(POPPLER数据集) 本仓库包含以下内容: 1. Code.zip:包含可复现本研究部分内容的代码: a. 1_generate_datasets:实现我们的数据集生成方法论 b. 2_run_models:在评估流程中运行机器学习模型 c. 3_result_replication:基于机器学习评估结果生成论文中展示的图表 2. Datasets.zip:包含2个文件夹: a. 原始数据集:1份源自NVD Vuldeepecker,3份从BigVul数据集提取得到 b. POPPLER数据集:基于我们的方法论从BigVul数据集中提取的、针对poppler项目的各时间观测点对应的训练集、验证集与测试集 3. Pretrained-models.zip:我们在评估阶段生成的预训练模型(对应时间线[2009-2018]中每个时间点的3组测试结果) 4. Results.zip:本研究的评估结果,其中ALL文件夹包含整体评估结果,其余文件夹为按模型划分的细分评估结果 文档资料 - INSTALL.pdf:代码安装指南 - README.pdf:项目说明文档 - REQUIREMENTS.pdf:硬件与软件环境要求说明 - STATUS.pdf:工件提交状态说明 - LICENSE.pdf:本工件的许可证文件 - PAPER.pdf:论文的定稿(最终正式版) 其他数据集与预训练模型可参考以下仓库: - 第一部分 NVD Vuldeepecker:https://doi.org/10.5281/zenodo.8207883 - 第二部分 LINUX:https://doi.org/10.5281/zenodo.10960662 - 第三部分 OPENSSL:https://doi.org/10.5281/zenodo.10966117



