chembed training dataset
收藏资源简介:
This dataset is a filtered subset of PubChem molecules used to train our chembed main model. The original dataset was downloaded from https://ftp.ncbi.nlm.nih.gov/pubchem/Compound/Extras/CID-SMILES.gz on 03/11/2023. Mixtures were split, salts and isotope values were removed, compounds without C atoms or with atoms outside the organic atoms set (B, Br, C, Cl, F, H, I, N, O, P, S, Si, Sn) were removed, compounds with weight higher than 600 Da were removed. SMILES were standardized with the ChEMBL standardization pipeline (Bento et al., 2020) and RDKit. Duplicates were removed. SELFIES were added using the selfies package (Krenn et al., 2022). Molecular properties (MolWt, MolLogP, TPSA, BertzCT, Kappa1, Kappa2, Kappa3) were computed with RDKit. The dataset was randomly split into train (80%) and test (20%) -> train.parquet, test.parquet.Please cite:- Kim, Sunghwan, et al. "PubChem 2023 update." Nucleic acids research 51.D1 (2023): D1373-D1380.- Bento, A. Patrícia, et al. "An open source chemical structure curation pipeline using RDKit." Journal of Cheminformatics 12.1 (2020): 51.- Krenn, Mario, et al. "SELFIES and the future of molecular string representations." Patterns 3.10 (2022).



