CompRep: A Dataset For Computational Reproducibility
收藏资源简介:
Reproducibility in computational science is increasingly dependent on the ability to faithfully re-execute experiments involving code, data, and software environments. However, assessing the effectiveness of reproducibility tools is difficult due to the lack of standardized benchmarks. To address this, we collected 38 computational experiments from diverse scientific domains and attempted to reproduce each using 8 different reproducibility tools. From this initial pool, we identified 18 experiments that could be successfully reproduced using at least one tool. These experiments form our curated benchmark dataset, which we release along with reproducibility packages to support ongoing evaluation efforts.
计算科学中的可复现性(Reproducibility)愈发依赖于对包含代码、数据与软件环境的实验开展精准重执行的能力。然而,由于缺乏标准化基准测试集,评估可复现性工具的有效性颇具挑战。为解决这一问题,我们从覆盖多个科学领域的研究中收集了38项计算实验,并尝试使用8种不同的可复现性工具对每一项实验进行重执行。从该初始实验池中,我们筛选出18项可通过至少一种工具成功复现的实验。这些实验构成了我们精心整理的基准数据集,我们将其与配套的可复现性工具包一同发布,以支持后续的评估工作。



