CompRep: A Dataset For Computational Reproducibility
收藏资源简介:
Reproducibility in computational science is increasingly dependent on the ability to faithfully re-execute experiments involving code, data, and software environments. However, assessing the effectiveness of reproducibility tools is difficult due to the lack of standardized benchmarks. To address this, we collected 38 computational experiments from diverse scientific domains and attempted to reproduce each using 8 different reproducibility tools. From this initial pool, we identified 18 experiments that could be successfully reproduced using at least one tool. These experiments form our curated benchmark dataset, which we release along with reproducibility packages to support ongoing evaluation efforts.
计算科学领域的可复现性(Reproducibility)愈发依赖于对包含代码、数据与软件环境的实验进行精准重执行的能力。然而,由于缺乏标准化基准测试集,评估可复现性工具的有效性颇具挑战。为解决这一问题,我们从多个不同科学领域收集了38项计算实验,并尝试使用8种不同的可复现性工具分别对每项实验开展复现操作。从初始候选集合中,我们筛选出可通过至少一种工具成功复现的18项实验,这些实验构成了我们精心整理的基准数据集。我们将该数据集与配套的可复现性工具包一并发布,以支持后续的评估研究工作。



