AD-MTL-Bench: A Multi-Task Benchmark Dataset for Alzheimer's Disease Drug Discovery
收藏资源简介:
AD-MTL-Bench v2 is a curated benchmark dataset for multi-task molecular property prediction in Alzheimer's disease and CNS drug discovery. This version expands the original 12-task dataset to 28 binary classification tasks across 157,163 unique compounds sourced from ChEMBL 37 and TDC ADMET benchmarks. Tasks span four axes: 13 efficacy targets (tau aggregation, Abeta aggregation, SIGMAR1, KIT, PDGFRA, PDGFRB, TMEM97, CETP, JAK1, JAK2, KDM1A, BRD4, MAPK1), 5 CNS safety endpoints (hERG, CYP3A4, CYP2D6, CYP2C9, Pgp), 2 BBB filter tasks, and 8 annotation targets including AD (AChE, BChE, GSK3B, MAOB), tau kinases (DYRK1A, CDK5), and Parkinson's disease/neuroinflammation (LRRK2, NLRP3). The dataset uses a Murcko scaffold split (125,758 / 18,031 / 12,721 train/val/test). Baseline: Uni-Mol MTL macro AUROC 0.946 +/- 0.002 (4 seeds, 28 tasks, scaffold test set).



