CompRep: A Dataset For Computational Reproducibility
收藏资源简介:
Reproducibility in computational science is increasingly dependent on the ability to faithfully re-execute experiments involving code, data, and software environments. However, assessing the effectiveness of reproducibility tools is difficult due to the lack of standardized benchmarks. To address this, we collected 38 computational experiments from diverse scientific domains and attempted to reproduce each using 8 different reproducibility tools. From this initial pool, we identified 18 experiments that could be successfully reproduced using at least one tool. These experiments form our curated benchmark dataset, which we release along with reproducibility packages to support ongoing evaluation efforts.
计算科学领域的可复现性(Reproducibility)愈发依赖于对包含代码、数据与软件环境的实验开展精准重执行的能力。然而,由于缺乏标准化基准(standardized benchmarks),评估可复现性工具的有效性极具挑战性。为解决这一问题,我们从多个不同科学领域收集了38项计算实验,并尝试使用8种不同的可复现性工具对每一项实验进行复现。从该初始实验集合中,我们筛选出18项可通过至少一种工具成功复现的实验。这些实验构成了我们经精心筛选的基准数据集,我们同步发布了配套的可复现性工具包,以支持相关评估工作的开展。



