遇见数据集

Leakage-Aware Evaluation of Molecular and Graph-Regularized Features for Cold-Start Drug-Disease Association Prediction (code and data)

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

GiG-MDA: Guilt-by-Association, Graph-Regularized and Molecular Features for Cold-Start Drug–Disease Association Ranking under a Leakage-Aware Protocol Version 9.24 — code and data accompanying the manuscript submitted to International Journal of Molecular Sciences (MDPI), Special Issue "Machine Learning Applications in Bioinformatics and Biomedicine: 4th Edition". Overview This archive contains the complete, versioned implementation and benchmark artifacts of the GiG-MDA leakage-aware evaluation pipeline: pair-disjoint and simulated zero-association cold-drug protocols on C-Dataset, F-Dataset, and DDCD, with all features, embeddings, negatives, and model selection constructed within the training partition. Key contents: Code (code/, ~40 Python scripts): data splitting and audit (data_split.py, audit_split.py, cold_split.py, scaffold_split.py), leak-safe feature construction (build_mirage_features.py, incl. the CTD-therapeutic full neighbor source and --exclude MolFormer protocol mode), committee-consensus negative filtering (negative_mining_oof.py, MolFormer-free by assertion), GRMF embeddings (pretrain_gigs_split.py, gigs_model.py), evaluation (eval_holdout_protocol.py, cold_eval.py, run_multiseed*.py), random-feature controls with 20 drug-level repeats (compare_pretrain_ablation.py), regularization diagnostics (regularization_diagnostic.py), matched model-selection pilot (matched_selection_pilot.py), the therapeutic-case-study pipeline (case_study_deploy.py + Crossref evidence retrieval and independent second-reviewer files), figure generation, and two scripts supporting the cold-start table end to end: aggregate_cold_results.py (regenerates the cold-start result table from the training-local pipeline) and dump_cold_predictions.py (writes per-pair predictions and recomputes AUROC / average precision). Results (results/): per-seed ledgers and all manuscript tables (regular, cold, cold-disease, scaffold, pretraining controls with 20 draws, calibration/top-k, negative-sampling comparison, GCN probe, regularization diagnostic, case-study tables with per-pair literature evidence and adjudication notes). Data (data/): public benchmark datasets C-Dataset/F-Dataset/DDCD as originally released, split manifests for all reported dataset/seed combinations, committee-filtered training negatives, and the CTD-derived therapeutic-only subset (DDCD/Mapping/mapping_therapeutic.csv, 13,579 pairs) with its name-matching audit. Reproduction record (reproduction/): per-pair predictions (candidate identifiers, labels, and predicted scores) for the four main XGBoost configurations on the C-Dataset cold-drug splits, produced with device=cuda and with device=cpu, together with per-split summaries and a README documenting the comparison. Under device=cuda the released code reproduces the reported AUPR for all four configurations in both splits; under device=cpu the same inputs give different values, and the seed-42 split reverses the order of the GBA baseline and the three-channel configuration. MoLFormer artifacts (code/results/molformer/): frozen 768-d SMILES embeddings and similarity matrices used by the molecular channel and controls. Figures (figures/), documentation (docs/), requirements.txt, and a pinned environment-lock.txt. Reproducibility Environment: environment-lock.txt pins the exact versions used for the reported values (Python 3.12.7, XGBoost 3.0.2, scikit-learn 1.7.2, RDKit 2026.03.5, NumPy 1.26.4, pandas 2.2.2); requirements.txt retains the original version ranges. All reported values were produced with device='cuda'. XGBoost's histogram method is not bit-identical between its GPU and CPU implementations, so closely spaced configurations can change order across devices; the measured effect is documented in reproduction/README.md and in the manuscript. Canonical entry points and the exact command sequences for each manuscript table are listed in README.md (Sections 3–4) and in the per-artifact README notes. The largest DDCD evaluation files exceed per-file limits and remain on the original Zenodo record (DOI: 10.5281/zenodo.21883418); restore them with code/download_ddcd_zenodo.py. SHA-256 of this archive is provided in the companion .sha256 file. License and data-use notes Code: MIT license (see repository). DrugBank-derived source attributes (category, condition, description, mechanism, pharmacodynamics, SMILES, targets) are not redistributed in this archive; users must obtain the applicable DrugBank release under its license and run the supplied preprocessing scripts. CTD data are redistributed under CTD's terms (comparative toxicogenomics database, ctdbase.org). Version history 9.24: added the reproduction record (reproduction/, per-pair predictions under CUDA and CPU, with the measured device effect), aggregate_cold_results.py and dump_cold_predictions.py, and the pinned environment-lock.txt; corrected two stale entries in the cold-start table (C-Dataset seeds 42 and 7, three-channel configuration) that had not been updated after the negative-mining isolation fix; updated author list, funding acknowledgements, and journal target (International Journal of Molecular Sciences). 9.3: therapeutic case study (deployment model + CTD therapeutic-only known set, three-view literature evidence with independent double review), P0-1 negative-mining isolation (committee features exclude pretrained-similarity columns; verified zero change over all 23 evaluated splits), conservative abstract/conclusions, BibTeX references. Earlier revisions archived on GitHub (https://github.com/ljx999h/GiG-MDA). Please cite Li, W.; Hu, Q.; Fu, Y.; Zhang, X.; Liu, B.; Mai, T.; Wang, P. "GiG-MDA: Guilt-by-Association, Graph-Regularized and Molecular Features for Cold-Start Drug–Disease Association Ranking under a Leakage-Aware Protocol" (manuscript under review). For the benchmark protocol details, see also the README and the code comments.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务