Training split MS2Deepscore
收藏资源简介:
The training data split into training, validation and test set used for training the ms2deepscore model found at: https://doi.org/10.5281/zenodo.13897744The spectra are a combination of GNPS spectra, MassBank, Mona and a library created by Corinna Brungs (see MSnLib Mass spectral libraries (.mgf and .json) (zenodo.org)) preprossessed using matchms to harmonize metadata. If you use this for academic work, please make sure to cite all relevant work.The data was split by selecting 1/20th of the unique inchikeys for the training set and the validation set. So no compound in the training data appears in the test or train set. For the rest the split is random.
本训练数据已划分为训练集、验证集与测试集,用于训练ms2deepscore模型,相关模型资源可访问:https://doi.org/10.5281/zenodo.13897744。该质谱谱图数据集整合了GNPS谱图、MassBank、Mona以及Corinna Brungs构建的谱库(详见MSnLib质谱谱库(.mgf与.json格式) (zenodo.org)),并通过matchms工具完成预处理以统一元数据格式。若将本数据集用于学术研究,请务必引用所有相关文献。本数据集的划分规则为:从唯一InChIKey (International Chemical Identifier Key)中选取1/20作为训练集与验证集的化合物,因此训练数据中的所有化合物均不会出现在测试集内,剩余部分则采用随机划分方式。



