遇见数据集

SyntheMol-RL Data

收藏
Zenodo2026-04-23 更新2026-05-26 收录
官方服务:

资源简介:

This archive contains data accompanying the paper "SyntheMol-RL: a reinforcement learning framework for designing novel and synthesizable antibiotics". The data is structured as follows. property_prediction_models: Chemprop, Chemprop-RDKit, and MLP-RDKit models trained to predict S. aureus growth inhibition or aqueous solubility. For each model, dataset, and dataset split (cross-validation by default or scaffold split), an ensemble of ten models was trained. supplementary_tables: The supplementary tables accompanying the paper with data such as the training data for the property prediction models and the compounds that were generated, synthesized, and tested. synthemol_generations: Molecules generated by SyntheMol (RL-Chemprop, RL-MLP, or MCTS versions) across different settings (ablation experiments) and random seeds (five per setting). synthesized_spectra: The spectra (LC-MS or 1H-NMR) of all synthesized compounds to verify their purity, split by synthesis vendor (Enamine or WuXi). The IDs used to label the spectra correspond to the IDs in the "id" column of the sheet named "Ordered Molecules" in Table_S11.xlsx in the supplementary_tables folder. vs_chemprop: The full set of 21 million molecules (14 million from Enamine REAL, 7 million from WuXi GalaXi) with Chemprop-RDKit property predictions used in the virtual screening (VS-Chemprop) baseline.

提供机构:
Zenodo
创建时间:
2025-05-12
二维码
社区交流群
二维码
科研交流群
商业服务