Reported model performance of past publications.
收藏资源简介:
The mapping of metabolite-specific data to pathways within cellular metabolism is a major data analysis step needed for biochemical interpretation. A variety of machine learning approaches, particularly deep learning approaches, have been used to predict these metabolite-to-pathway mappings, utilizing a training dataset of known metabolite-to-pathway mappings. A few such training datasets have been derived from the Kyoto Encyclopedia of Genes and Genomes (KEGG). However, several prior published machine learning approaches utilized an erroneous KEGG-derived training dataset that used SMILES molecular representations strings (KEGG-SMILES dataset) and contained a sizable proportion (~26%) duplicate entries. The presence of so many duplicates taint the training and testing sets generated from k-fold cross-validation of the KEGG-SMILES dataset. Therefore, the k-fold cross-validation performance of the resulting machine learning models was grossly inflated by the erroneous presence of these duplicate entries. Here we describe and evaluate the KEGG-SMILES dataset so that others may avoid using it. We also identify the prior publications that utilized this erroneous KEGG-SMILES dataset so their machine learning results can be properly and critically evaluated. In addition, we demonstrate the reduction of model k-fold cross-validation (CV) performance after de-duplicating the KEGG-SMILES dataset. This is a cautionary tale about properly vetting prior published benchmark datasets before using them in machine learning approaches. We hope others will avoid similar mistakes.
将代谢物特异性数据映射至细胞代谢通路,是生化阐释过程中至关重要的数据分析步骤。诸多机器学习方法,尤其是深度学习方法,已被用于预测此类代谢物-通路映射关系,其训练数据来自已知的代谢物-通路映射数据集。此类训练数据集中有少数源自京都基因与基因组百科全书(Kyoto Encyclopedia of Genes and Genomes, KEGG)。然而,此前已有多项机器学习研究使用了一套存在错误的KEGG衍生训练数据集——该数据集采用SMILES(Simplified Molecular-Input Line-Entry System)分子表征字符串,且存在占比约26%的大量重复条目。如此大量的重复条目会污染基于KEGG-SMILES数据集开展k折交叉验证所生成的训练集与测试集。因此,由于这些重复条目的错误存在,所得机器学习模型的k折交叉验证性能被严重高估。本文对KEGG-SMILES数据集进行了描述与评估,以期其他研究者避免误用该数据集。同时,本文还梳理了所有使用过这套错误KEGG-SMILES数据集的已发表研究,以便学界能够对相关机器学习研究的结果进行严谨且批判性的评估。此外,本文还验证了对KEGG-SMILES数据集去重后,模型的k折交叉验证性能会出现下降。本研究为学界敲响警钟:在机器学习研究中使用已发表的基准数据集前,应当对其进行严格的审核与验证。我们期望其他研究者能够避免类似的错误。




