Small-Molecule Anticancer Compounds for Breast Cancer Drug Discovery
收藏资源简介:
This dataset comprises small-molecule anticancer compounds collected from publicly available datasets and previous studies on anticancer drug discovery. The primary sources are the training set reported by Al-Jarf et al. and the ChEMBL database. The initial merged dataset contains 116,008 molecules, each represented by its Simplified Molecular Input Line Entry System (SMILES) string. The SMILES representations were processed to remove duplicate entries and validate their chemical syntax. To construct the active anti-breast cancer dataset, compounds were collected from six breast cancer cell lines: BT549, HS_578T, MCF7, MDA_MB_231_ATCC, MDA_MB_468, and T47D. The corresponding GI50 values were extracted and used to determine compound activity. Following the removal of duplicate entries, compounds with GI50 values less than or equal to 5 were classified as active. This filtering procedure resulted in a subset of 43,551 active compounds, which was subsequently used for model training and molecular generation. The dataset supports the training of the Variational Autoencoder (VAE) to learn a continuous latent representation of the chemical space. The resulting active compound dataset also serves as the basis for the genetic algorithm (GA)-based generation and optimization of candidate anticancer molecules.



