遇见数据集

Replication Package for Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects"

收藏
Zenodo2022-09-08 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models. We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage. We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars. The raw data of repositories are stored in `Replication Package/data/prodata.pkl`, and the contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m2_k1.pkl`. We sort items by the order in `Replication Package/data/randind.npy`, which can be used to reproduce the same results as in the paper. <br> `Replication Package/data/X_test_m2_k1.pkl` and `Replication Package/data/y_test_m2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table III and IV, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table II. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section III.B.

本复现包可用于复现本论文中的全部实验结果。其包含两部分内容:1)290255个代码仓库的数据集;2)用于模型训练与解释的Python脚本。 我们建议在通用Linux主机上手动配置所需运行环境,该主机需至少配备1个CPU核心、8GB内存及100GB可用存储空间。本次开发与所有实验均在搭载两颗英特尔至强金牌CPU、320GB内存与36TB RAID 5存储阵列的Ubuntu 20.04服务器上完成。 我们通过GHTorrent还原了290255个代码仓库的历史状态,这些仓库均满足:代码提交次数超过57次、拉取请求(Pull Request,PR)≥4个、议题(Issue)≥1个、复刻(Fork)≥1次、星标(Star)≥2个。 代码仓库的原始数据存储于`Replication Package/data/prodata.pkl`中,而局部可解释模型无关解释(Local Interpretable Model-agnostic Explanations,LIME)模型生成的特征贡献度结果存储于`Replication Package/data/limeres_m2_k1.pkl`。我们根据`Replication Package/data/randind.npy`中记录的顺序对数据项进行排序,该文件可用于复现与论文完全一致的实验结果。 <br> `Replication Package/data/X_test_m2_k1.pkl`与`Replication Package/data/y_test_m2_k1.pkl`存储了用于LIME模型的测试数据集。您可通过运行`Replication Package/fitdata.py`获取论文中表III与表IV的结果,运行`Replication Package/draw_compare_variable.py`生成图2,运行`Replication Package/allvari_statistics.py`获取表II的结果。在`Replication Package/Variable_comparison_with_different_parameter.pdf`中,我们展示了不同参数设置下的LIME模型实验结果。此外,`Replication Package/sample_pros.csv`还提供了论文第三章B节中随机选取的代码仓库列表。

提供机构:
Zenodo
创建时间:
2022-09-08
二维码
社区交流群
二维码
科研交流群
商业服务