遇见数据集

AMRK-DB: an explainable, multi-organism, lineage-validated AMR biomarker knowledge base for ESKAPEE pathogens (45 models, 6 organisms, 14 antibiotic classes)

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

AMRK-DB is an explainable, multi-organism, FAIR, unitig-resolution knowledge base of machine-learning-derived genomic biomarkers of antimicrobial resistance in ESKAPEE pathogens. This deposit contains 45 models spanning 6 organisms (Escherichia coli, Klebsiella pneumoniae, Staphylococcus aureus, Acinetobacter baumannii, Pseudomonas aeruginosa, Enterococcus faecium) and 14 antibiotic classes, built from 78,556 genome-phenotype pairs obtained from BV-BRC. Features are unitigs from a compacted de Bruijn graph (no reference genome). Every model is validated with lineage-aware cross-validation (PopPUNK strain clusters as StratifiedGroupKFold groups); the reported mean ROC-AUC is 0.842 (range 0.429-0.975). A matched lineage-blind comparison, holding folds, stratification and hyperparameters fixed, shows that removing the lineage grouping inflates the AUC in all 45 models (mean +0.088, max +0.437; Wilcoxon p = 5.7e-14), and the inflation scales with clonal dominance. Each of the 3,571 biomarkers carries seven orthogonal evidence layers - CARD/NCBI BLAST tiering with ARO ontology mapping, resistant-versus-susceptible prevalence, CARD variant-model SNP allele checks, MDA and label-permutation nulls, CPSS stability selection with a PFER bound, and pyseer LMM population-structure-corrected significance - combined into one evidence tier per (biomarker, model): 349 confirmed, 942 candidate, 1,920 weak, 337 none and 23 strong_novel. The strong_novel tier is the contribution a BLAST-only view hides: stable, LMM-significant biomarkers with no known resistance gene. Contents: the unified SQLite knowledge base (amrk.db, schema 0.7.1), five tidy tables (model summary, KB overview, full biomarker evidence, mechanisms, random-versus-lineage CV, cross-antibiotic overlap), 36 figures, per-model candidate and evidence outputs, and run metadata recording the git commit, random seed, config hash and the version of every tool that shaped the result (unitig-caller, PopPUNK, graph-tool, BLAST, pyseer, CARD snapshot). Not included: unitig matrices and raw assemblies, which the pipeline regenerates from BV-BRC. Limitations: associations are statistical, not mechanistic. Labels are BV-BRC's curated phenotypes (raw MIC re-interpretation was not feasible - fill rate is 9% for S. aureus). External validation is lineage hold-out plus AMRFinderPlus/ResFinder concordance; temporal validation is impossible (BV-BRC AMR data ends in 2021) and geographic hold-out is confounded by country dominance. Some models exceed a PFER of 1, and co-carried genes (e.g. class-1 integron markers) reflect linkage, not causation. Software: https://github.com/demirbase/ML_AMR_Prediction_v2 (MIT). Data: CC-BY-4.0.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务