遇见数据集

AD-MTL-Bench: A Multi-Task Benchmark Dataset for Alzheimer's Disease Drug Discovery

收藏
Zenodo2026-05-17 更新2026-05-26 收录
官方服务:

资源简介:

AD-MTL-Bench is a curated benchmark dataset for multi-task molecular property prediction in Alzheimer's disease (AD) drug discovery. It integrates 63,105 unique compounds across 12 binary classification tasks spanning three pharmacological axes: AD disease-modifying targets (BACE1, GSK3β, MAO-B), symptomatic cholinesterase targets (AChE, BChE), and CNS ADMET liability endpoints (BBB permeability from two sources, hERG, CYP3A4, CYP2D6, CYP2C9, P-glycoprotein). The dataset is split into train/validation/test (49,832 / 6,310 / 6,310) using a Murcko scaffold-based split that ensures no scaffold appears in more than one partition, providing a realistic evaluation of generalization to structurally novel compounds. Each compound record includes: InChIKey, RDKit-standardized SMILES, Murcko scaffold SMILES, PAINS and Brenk structural alert flags, binary task labels (with NaN for unlabeled entries), and split assignment. Baseline AUROC on the scaffold test set: Chemprop single-task 0.835, Chemprop multi-task 0.864, Uni-Mol multi-task 0.893 (macro mean across all 12 tasks, 5 seeds). Data sources: ChEMBL 33 (BACE1, AChE, BChE, GSK3β, MAO-B, hERG), B3DB (BBB permeability), and TDC ADMET Benchmark (CYP3A4, CYP2D6, CYP2C9, Pgp, BBB-Martins). Files included: - multitask_dataset.csv: full compound table with all labels - split_scaffold.csv: scaffold split assignments (inchikey, split) - dataset_card.md: full documentation (schema, statistics, baselines, known biases)

提供机构:
Zenodo
创建时间:
2026-05-17
二维码
社区交流群
二维码
科研交流群
商业服务