遇见数据集

chembed training dataset

收藏
Zenodo2025-10-06 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is a filtered subset of PubChem molecules used to train our chembed main model. The original dataset was downloaded from https://ftp.ncbi.nlm.nih.gov/pubchem/Compound/Extras/CID-SMILES.gz on 03/11/2023. Mixtures were split, salts and isotope values were removed, compounds without C atoms or with atoms outside the organic atoms set (B, Br, C, Cl, F, H, I, N, O, P, S, Si, Sn) were removed, compounds with weight higher than 600 Da were removed. SMILES were standardized with the ChEMBL standardization pipeline (Bento et al., 2020) and RDKit. Duplicates were removed. SELFIES were added using the selfies package (Krenn et al., 2022). Molecular properties (MolWt, MolLogP, TPSA, BertzCT, Kappa1, Kappa2, Kappa3) were computed with RDKit. The dataset was randomly split into train (80%) and test (20%) -> train.parquet, test.parquet.Please cite:- Kim, Sunghwan, et al. "PubChem 2023 update." Nucleic acids research 51.D1 (2023): D1373-D1380.- Bento, A. Patrícia, et al. "An open source chemical structure curation pipeline using RDKit." Journal of Cheminformatics 12.1 (2020): 51.- Krenn, Mario, et al. "SELFIES and the future of molecular string representations." Patterns 3.10 (2022).

提供机构:
Zenodo
创建时间:
2025-10-06
二维码
社区交流群
二维码
科研交流群
商业服务