遇见数据集

Molecular fingerprint benchmarking datasets

收藏
Zenodo2026-02-18 更新2026-05-26 收录
官方服务:

资源简介:

This is a collection of several datasets used to benchmark molecular fingerprint algorithms.ms2structures dataset (37,811 compounds)We assembled a curated collection of compounds measured and annotated by tandem mass spectrometry. We merged the training and evaluation data from MS2Deepscore, a deep-learning model for predicting chemical similarity from mass spectra with the overlapping benchmarking set, MassSpecGym. biostructures dataset (718,067 compounds)Starting from 718,097 biologically relevant compounds drawn from the dataset used by Kretschmer et al. [2025], we removed 30 entries that RDKit could not convert to fingerprints. Using the Classyfire API we added chemical class information for 695,152 compounds. 25-subclasses dataset (75,000 compounds) Of 25 of the 27 most common chemical subclasses, excluding x and y since they were too easily distinguishable from the rest, 3000 compounds were randomly sampled from the biostructures dataset. This results in a balanced classification dataset. 120-subclasses dataset (120,000 compounds) This is another classification-task-oriented dataset, now with a random sample of 1000 unique compounds for the 120 most common chemical subclasses from the biostructures dataset. rascalMCES dataset (5,413,677 compound pairs)We randomly sampled 5,557,963 compound pairs from the ms2structures dataset whose precursor masses differ by at most 100 Da. RascalMCES scores were computed with RDKit on an Intel Core i9-13900K (settings: similarityThreshold=0.05, maxBondMatchPairs=1000, minFragSize=3, timeout=60 s). After excluding 144,286 timed-out pairs, our final benchmarking set contained 5,413,677 pairs

提供机构:
Zenodo
创建时间:
2026-02-18
二维码
社区交流群
二维码
科研交流群
商业服务