遇见数据集

Pharmacological Lattice Quantisation

收藏
Zenodo2026-07-05 更新2026-08-02 收录
官方服务:

资源简介:

An honest computational platform for neglected-disease drug discovery: ligand-based triage + structure-based generative design (antimicrobial resistance & malaria) Type: Software + data + computational-prediction dataset (research prototype) > READ FIRST. This dataset has two layers. The ligand-based layer is validated on held-out > (scaffold-split) data and reports real, reproducible metrics. The structure-based (SBDD) layer > is positive-control-validated as a method, but the molecules it generates are computational > predictions with no experimental validation and no benchmarking, hypotheses to test, not hits. > Neither layer models clinical efficacy. For research prioritisation, not clinical use. Description This dataset supports a two-layer computational platform for neglected-disease drug discovery, focused on antimicrobial resistance and malaria: (A) a ligand-based drug-triage and de-risking platform, and (B) a structure-based generative design pipeline (SBDD) built on top of it. It provides cleaned, reproducible, honestly-evaluated bioactivity data together with the results needed to judge model reliability, plus the full structure-based pipeline and the candidate molecules it produced. A. Ligand-based triage and de-risking platform Data were derived from two open resources: ChEMBL (bioactivity, CC BY-SA 3.0) and the Open Targets Platform (human target-disease genetics, CC BY 4.0). Bioactivity records were standardised to a p-scale (p = -log10 of the molar concentration; pIC50 for binding assays, pMIC for whole-cell antibacterial assays, converting ug/mL via molecular weight), deduplicated by InChIKey with replicate measurements aggregated by median, and canonicalised with RDKit. Each compound carries a Bemis-Murcko scaffold split (train/test) so the held-out evaluation can be reproduced exactly, with no analog leakage between splits. This layer contains: (1) five bioactivity tables - Plasmodium falciparum (malaria, IC50), Staphylococcus aureus and Escherichia coli (whole-cell MIC), acetylcholinesterase (AChE, IC50), and EGFR (IC50, oncology reference); (2) a hERG cardiotoxicity liability table, where censored "greater-than" records are reclaimed as confident non-blockers - the valuable negative data usually discarded; (3) Open Targets target-validation evidence, including protective loss-of-function (LoF) signals; (4) held-out model metrics and figures; and (5) the full code module that regenerates everything deterministically (2048-bit ECFP4 count fingerprints, RandomForest, fixed seed). Notable findings. Held-out (scaffold-split) ranking performance spans Spearman 0.51-0.77; whole-cell MIC is intentionally harder to predict than single-target binding because it folds in membrane permeability and efflux. Prediction reliability is quantified by RandomForest tree-variance: withholding the least-confident 40% of malaria predictions raises Spearman from ~0.55 to ~0.68 (a validated abstention / applicability-domain effect). The hERG classifier reaches ROC-AUC 0.80. Genetic evidence independently confirms target relevance: AChE->Alzheimer's disease carries no genetic association (a symptomatic target), whereas PCSK9 shows protective LoF variants (odds ratios ~0.35-0.61) - a natural-knockout, prevention-oriented signal. B. Structure-based generative design pipeline (SBDD) Building on the ligand layer, a pathogen-agnostic structure-based pipeline discovers essential-but- undrugged protein targets and designs candidate binders for them. For a chosen pathogen it: (1) finds targets that are **essential** (real gene-essentiality data - BV-BRC for bacteria; PlasmoDB piggyBac Mutagenesis Index Scores for Plasmodium), undrugged, host-selective (no close human homolog), and druggable (site-anchored pocket hydrophobicity/enclosure), with resistance barriers from CARD; (2) fetches the AlphaFold structure and UniProt catalytic residues; (3) scores small-molecule binding with a positive-control-validated Boltz-2 co-folding oracle returning a binding probability and a 3D pose - validated per target against known drug/target pairs (e.g. DHFR/methotrexate 0.998 vs decoy 0.13; A. baumannii FabI/triclosan lifted 0.577->0.912 in the FabI+NAD+ ternary complex), where rigid docking (AutoDock Vina) failed the same control and the oracle is shown deterministic (sigma ~0.003); (4) checks whether the co-folded pose engages the catalytic site; (5) generates candidate molecules with a Markov-chain / fragment-crossover / substrate-growing generator **seeded from each target's native substrate or inhibitor pharmacophore**, under a baked-in reactive-group/PAINS/charge/catechol/ thiocarbonyl filter that prevents oracle-gaming (documented and rejected artifacts: boron and thiolate exploits); and (6) diagnoses the binding ceiling as a search, oracle, target, or cofactor limit (e.g. the FabI ceiling was pre-registered as cofactor-limited and confirmed by the NAD+ lift). Demonstrated cross-kingdom on Mycobacterium tuberculosis (bacterial), Plasmodium falciparum (including PfDHODH, a resistance-independent, host-selective target relevant to artemisinin (K13)-resistant malaria**), and the WHO-critical pan-resistant ESKAPE pathogen Acinetobacter baumannii** (FabI). This layer includes the full pipeline code, the SQLite results store, a lab-ready combination dossier, and candidate-molecule SDFs with honest provenance (e.g. novel biaryl-pyridinone A. baumannii FabI leads). Every SDF record is a computational prediction to be tested, not a hit. Honesty and scope (applies to both layers) The ligand models are triage/ranking tools with an explicit applicability domain: they should abstain on compounds far from their training chemistry rather than extrapolate. The SBDD layer has no experimental validation, no benchmarking against standard SBDD sets, and mostly modest predicted affinities; the per-target enzyme activity assay is the confirmation gate that would turn a predicted binder into a confirmed inhibitor. Crucially, neither layer models whether engaging a target treats disease in humans (clinical efficacy), which lies outside ligand structure and target binding and is deliberately not included. This is decision-support data for early-stage prioritisation and de-risking. How to interpret and use In bioactivity tables, higher `p_affinity` means more potent binding/inhibition; use the `scaffold_split` column to reproduce the reported metrics. In the SBDD SDFs, a higher predicted binding probability means a stronger *predicted* interaction, but these are unvalidated computational designs - treat them as ranked hypotheses and run the target's enzyme assay to confirm. See `docs/REPLICATION.md` for the exact environment, data queries, limitations, and known bugs. Data sources (external, re-fetchable - not redistributed here) ChEMBL (bioactivity) · Open Targets Platform (target-disease genetics) · BV-BRC (bacterial essentiality) · PlasmoDB / Figshare Pf3D7_gene_annotations (Zhang 2018 piggyBac MIS) · CARD (resistance) · AlphaFold DB + UniProt (structures).

提供机构:
Zenodo
创建时间:
2026-07-05
二维码
社区交流群
二维码
科研交流群
商业服务