遇见数据集

Large scale RDKit based filtering of 123.9 million compounds followed by ADMET-AI profiling of 49 million candidates

收藏
Zenodo2026-06-16 更新2026-06-17 收录
官方服务:

资源简介:

The SMILES representations of compounds were retrieved from the COCONUT [1], LOTUS [2], NPASS [3], NPAtlas [4], and PubChem [5] databases on 18 April 2026. The initial dataset comprised 738,843 compounds from COCONUT, 276,518 from LOTUS, 203,390 from NPASS, 36,454 from NPAtlas, and 123,857,298 from PubChem, yielding a total of 125,112,503 SMILES. All SMILES were processed using RDKit (2025.09.6) (https://www.rdkit.org), which parsed and sanitized each SMILES string to ensure validity. During this step, 43,464 entries containing invalid SMILES syntax or valence errors were excluded. The remaining 125,069,039 valid molecules were converted into canonical isomeric SMILES and subsequently subjected to global deduplication. This process identified and removed 1,162,586 duplicate structures, resulting in a final, non-redundant library containing 123,906,453 unique canonical compounds. The dataset consisted of 738,823 compounds from COCONUT, 276,518 from LOTUS, 201,972 from NPASS, 36,453 from NPAtlas, and 122,652,687 from PubChem. This comprehensive and deduplicated library served as the starting point for all subsequent drug likeness filtering and ADMET profiling. To enrich the library for compounds with favorable drug like characteristics, the dataset was subjected to a sequential filtering workflow implemented in RDKit. The filtering workflow began with the combined application of Lipinski's Rule of Five and the Veber criteria, reducing the library from 123,906,453 to 86,482,856 compounds (69.8% retention). The remaining molecules were further filtered based on molecular complexity by retaining compounds with an Fsp³ value greater than 0.3 and no more than four aromatic rings, resulting in 57,681,058 compounds (66.7% retention). To eliminate molecules with a high likelihood of producing assay artifacts, the library was screened against the PAINS (Pan-Assay Interference Compounds) structural alerts, removing only a small fraction of compounds and retaining 56,225,456 molecules (97.5% retention). Finally, compounds were evaluated using the Quantitative Estimate of Drug-likeness (QED), and only those with QED scores greater than 0.5 were retained. This final step produced a high quality library comprising 49,068,327 compounds (87.3% retention). The resulting dataset contains 49,068,327 compounds and serves as the input library for comprehensive ADMET prediction using the standalone implementation of ADMET-AI version 2.0.1 [6]. All filtering and ADMET prediction scripts used to generate the dataset are available on GitHub: https://github.com/zubairbty/chemistry_admet-other References: [1] Chandrasekhar, Venkata, et al. "COCONUT 2.0: a comprehensive overhaul and curation of the collection of open natural products database." Nucleic Acids Research 53.D1 (2025): D634-D643. [2] Rutz, Adriano, et al. "The LOTUS initiative for open knowledge management in natural products research." elife 11 (2022): e70780. [3] Lin, Hanbo, et al. "NPASS database update 2026: comprehensive quantitative composition, bioactivity, and ADME-Tox data of natural products for biomedical research." Nucleic Acids Research 54.D1 (2026): D1519-D1527. [4] Poynton, Ella F., et al. "The Natural Products Atlas 3.0: extending the database of microbially derived natural products." Nucleic Acids Research 53.D1 (2025): D691-D699. [5] Kim, Sunghwan, et al. "PubChem 2025 update." Nucleic acids research 53.D1 (2025): D1516-D1525. [6] Swanson, Kyle, et al. "ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries." Bioinformatics 40.7 (2024): btae416.

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务