遇见数据集

ProCMT-QSAR: An Interpretable Target-Aware Protein-Conditioned Multimodal Multi-Task QSAR Model for Toxicity Prediction of Endocrine-Disrupting Chemicals

收藏
Zenodo2026-08-01 更新2026-08-01 收录
官方服务:

资源简介:

This repository provides the curated dataset and pretrained model checkpoints supporting ProCMT-QSAR, a protein-conditioned multimodal multi-task QSAR model for predicting endocrine-disrupting chemical (EDC) activity across six endocrine-relevant protein targets: androgen receptor (AR), aromatase (CYP19A1), estrogen receptor alpha (ESR1α), estrogen receptor beta (ESR2β), 11β-hydroxysteroid dehydrogenase type 1 (HSD11B1), and peroxisome proliferator-activated receptor gamma (PPARγ). Dataset. Bioactivity records for all six targets were retrieved from ChEMBL (AR: CHEMBL1871; CYP19A1: CHEMBL1978; ESR1: CHEMBL206; ESR2: CHEMBL242; HSD11B1: CHEMBL4235; PPARG: CHEMBL235), restricted to IC50 measurements on Homo sapiens targets with exact standard relation, assay confidence score ≥ 7, and complete mandatory fields (Molecule ChEMBL ID, SMILES, Standard Value, Standard Units). Corresponding canonical target sequences were retrieved from UniProt (P10275, P11511, P03372, Q92731, P28845, P37231) for protein embedding generation. Records were curated using a standardized pipeline (structure validation, salt stripping, canonicalization, duplicate aggregation by median IC50) and labeled under a clean-margin binary scheme (active: IC50 ≤ 1,000 nM; non-active: IC50 ≥ 10,000 nM; intermediate values excluded). Compounds were partitioned per endpoint using a stratified Bemis–Murcko scaffold split (80% train pool / 20% held-out test), verified to be free of scaffold and canonical-SMILES leakage between splits. This dataset is provided to enable full reproducibility of data curation, splitting, and downstream modeling steps described in the associated manuscript. Model checkpoints. This repository also provides the pretrained ProCMT-QSAR model artifacts required to reproduce or reuse the final trained model without retraining from scratch, including: Optuna hyperparameter search databases (10-fold cross-validation and medium-safe search configurations) Best hyperparameter configurations identified via Optuna (JSON) Final retrained model weights (PyTorch checkpoint) Fitted Mordred descriptor preprocessing object, required to transform new compounds consistently with the training pipeline Full source code for data preprocessing, molecular representation generation, model training, evaluation, and interpretability analysis is maintained separately at: https://github.com/kit-cml/ProCMT-QSAR

提供机构:
Zenodo
创建时间:
2026-08-01
二维码
社区交流群
二维码
科研交流群
商业服务