遇见数据集

Reproducibility levels.

收藏
Figshare2025-12-04 更新2026-04-28 收录
官方服务:

资源简介:

The BioModels Repository contains over 1000 manually curated mechanistic models from published literature, most often encoded in the Systems Biology Markup Language (SBML). This community-based standard formally specifies each model, but does not describe the computational experimental conditions to run a simulation and collect data. Therefore, it can be challenging to reproduce any figure or result from a publication with an SBML model alone. The Simulation Experiment Description Markup Language (SED-ML) provides a solution: a standard way to specify exactly how to run an experiment corresponding to a specific figure or result. BioModels was established years before SED-ML, and both systems evolved over time, both in content and acceptance. Hence, only about half of the entries in BioModels contained SED-ML files, and these files reflected the version of SED-ML that was available at the time. Additionally, almost all of these SED-ML files had at least one minor mistake that made them impossible to run. To make these models and their results more reproducible, we report here on our work updating, correcting and generating new SED-ML files for 1055 curated mechanistic models in BioModels. In addition, because SED-ML is implementation-independent, it can be used for verification, demonstrating that results hold across multiple simulation engines. We tested, corrected, and improved over 450 existing SED-ML files in the BioModels database, and created basic files for the rest of the entries. Then, we used a wrapper architecture for interpreting SED-ML, and report verification results across five different ODE-based biosimulation engines, after further improving the models, the wrappers, and the engines themselves. Our work with SED-ML and the BioModels collection aims to improve the utility of these models by making them more reproducible and credible. Improved reproducibility means these models are now even more fit for re-use, such as in new investigations and as components of multiscale models.

BioModels 数据库(BioModels Repository)收录了逾千份经人工整理的机制模型,这些模型均源自已发表的学术文献,其中绝大多数以系统生物学标记语言(Systems Biology Markup Language, SBML)进行编码。该基于社区的标准可对每一份模型进行形式化规范,但并未描述运行仿真并采集数据所需的计算实验条件。因此,仅依靠SBML模型,往往难以复现已发表论文中的任一图表或研究结果。仿真实验描述标记语言(Simulation Experiment Description Markup Language, SED-ML)则提供了解决方案:该标准可精准指定如何运行与特定图表或研究结果对应的实验。由于BioModels数据库的建立早于SED-ML问世多年,二者均随时间在内容与认可度层面不断演进,因此该数据库中仅约半数条目包含SED-ML文件,且这些文件均适配当时可用的SED-ML版本。此外,几乎所有现存的SED-ML文件都存在至少一处细微错误,导致其无法正常运行。为提升这些模型及其研究结果的可复现性,本文报道了我们的相关工作:为BioModels数据库中的1055份经整理的机制模型更新、修正并生成了全新的SED-ML文件。此外,由于SED-ML具备实现无关性,其可用于验证工作,以证明研究结果可在多种仿真引擎中复现。我们对BioModels数据库中现存的450余份SED-ML文件进行了测试、修正与优化,并为其余条目创建了基础SED-ML文件。随后,我们采用封装器架构对SED-ML进行解析,并在进一步优化模型、封装器及仿真引擎本身之后,报告了基于五种不同常微分方程(Ordinary Differential Equation, ODE)类生物仿真引擎的验证结果。我们针对SED-ML与BioModels数据集的相关工作,旨在提升这些模型的实用性,具体途径为增强其可复现性与可信度。可复现性的提升意味着这些模型如今更适用于二次研究,例如作为多尺度模型的组件或应用于全新的科研探索中。

创建时间:
2025-12-04
二维码
社区交流群
二维码
科研交流群
商业服务