Cross-ecosystem categorization: A manual-curation protocol for the categorization of Java Maven libraries along Python PyPI Topics (dataset)
收藏资源简介:
This dataset reports all information needed to implement a human-guided protocol for the categorisation of libraries, from any software ecosystem, along the 24 top-level PyPI Topic classifiers. It also contains the data produced in a demonstration, where the protocol was applied to 256 open-source Java libraries from Maven Central with high- or critical-severity CVEs. This dataset can be used as ground truth for cross-ecosystem studies in software engineering, especially from functional and security perspectives. This dataset contains: the protocol designed to interpret sources for category assessment, and arbitrate the results; the sources and metadata, including CVEs, collected for the demonstration; the set of categorised libraries and CVE statistics, including a higher-level classification into Local or Remote network functionalities.
本数据集收录了实现适用于任意软件生态系统中软件库分类的人工引导协议所需的全部信息,该协议基于24个顶级PyPI主题分类器(PyPI Topic classifiers)构建。本数据集还包含一项演示实验所生成的数据:该实验将上述协议应用于来自Maven中央仓库的256个存在高严重级或临界严重级通用漏洞与披露(Common Vulnerabilities and Exposures,以下简称CVE)的开源Java软件库。本数据集可作为软件工程领域跨生态系统研究的基准真值(ground truth),尤其适用于功能与安全视角下的相关研究。本数据集包含以下内容:1. 用于解读分类评估依据源数据并对结果进行仲裁的协议;2. 为本次演示实验收集的源数据与元数据(其中包含CVE);3. 已分类软件库集合与CVE统计数据,其中包含针对本地或远程网络功能的高阶分类。



