遇见数据集

A Cleaned Multi-Source MS/MS Dataset for Single-Modality Mass Spectral Retrieval and Cross-Instrument Evaluation

收藏
Zenodo2026-04-01 更新2026-05-26 收录
官方服务:

资源简介:

This benchmark provides a curated benchmark for single-modality tandem mass spectrometry (MS/MS) retrieval, designed to evaluate the generalization ability of machine learning models across different instruments and data sources. The benchmark integrates multiple public MS/MS datasets, including GNPS, MoNA, MassBank, MassSpecGym, MTBLS1572, and a large-scale simulated dataset (SimMS-9M). The spectra are organized by instrument types, primarily Orbitrap and QTOF. Data sources:- GNPS, MoNA, MassBank, MassSpecGym, and MTBLS1572 are derived from the dataset released in: "Supervised Contrastive Learning Leads to More Reasonable Spectral Embeddings" Paper: https://doi.org/10.1021/acs.analchem.5c02655 Data: https://doi.org/10.6084/m9.figshare.28876751 - SimMS-9M is obtained from: "CSU-MS2: A Contrastive Learning Framework for Cross-Modal Compound Identification from MS/MS Spectra to Molecular Structures" Paper: https://doi.org/10.1021/acs.analchem.5c01594 Data: https://zenodo.org/records/17756840 This dataset contains approximately 9 million simulated spectra generated from ~3 million compounds under three different collision energy settings. To construct a reliable and realistic benchmark, we adopt a strict non-overlapping protocol:- GNPS-Orbitrap is used exclusively for training.- All other datasets are used only for evaluation. To prevent data leakage, spectra corresponding to compounds present in the training set are removed from all evaluation datasets, ensuring no compound-level overlap. We do not redistribute the original datasets. Instead, we provide processed data, including filtered subsets, data splits, and evaluation files. This benchmark supports research in MS/MS retrieval, spectral representation learning, and cross-dataset generalization under realistic conditions.

提供机构:
Zenodo
创建时间:
2026-03-07
二维码
社区交流群
二维码
科研交流群
商业服务