Dataset for: An Integrated AI Framework for Targeted Depression Drug Discovery: Leveraging Cheminformatics and Genomics
收藏资源简介:
This dataset supports the findings of the manuscript titled: "An Integrated AI Framework for Targeted Depression Drug Discovery: Leveraging Cheminformatics and Genomics". This compressed archive (.zip) contains the key datasets generated and analysed during the study. The files include: Processed Bioactivity Data: The curated dataset derived from the ChEMBL database, including compound IDs, SMILES strings, and calculated pIC50 values used for model training. Molecular Descriptors Matrix: The final feature matrix containing the 1875 molecular descriptors computed using PaDEL-Descriptor for each compound. GWAS-derived Target List: The list of prioritized gene targets identified from the analysis of public Genome-Wide Association Study (GWAS) summary statistics. Final Candidate Compounds List: The prioritized list of candidate molecules linking potent compounds (predicted by the DNN model) to genetically relevant targets. The purpose of this data deposit is to ensure full reproducibility of our computational experiments and to provide a valuable resource for the scientific community. The primary raw data was sourced from ChEMBL and the Psychiatric Genomics Consortium (PGC).



