遇见数据集

HEG-TKG: Hierarchical Evidence-Grounded Temporal Knowledge Graphs for Rare Disease Reasoning - Paper 1 Reproducibility Bundle (v1.0.0)

收藏
Zenodo2026-05-09 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is the reproducibility bundle accompanying the manuscript "The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning" (Ahmed et al., submitted to npj Digital Medicine, 2026). The bundle covers every artifact required to reproduce the paper's quantitative claims, organised in three groups. Source data and graphs. 36 clinician-validated clinical scenarios across three phenotypically confusable rare neuromuscular disease pairs (Duchenne/Becker muscular dystrophy; myasthenia gravis / Lambert-Eaton myasthenic syndrome; CIDP / Guillain-Barré). Three temporal knowledge graphs with 5,481 nodes, 6,316 PMID-backed edges, 1,280 disease-trajectory milestones, and 987 unique source PMIDs, released as JSON, CSV, and Neo4j-compatible Cypher import scripts. 108 model outputs across the three evaluation arms (HEG-TKG, Vanilla LLM, Guideline-RAG), all using the same synthesis model (GPT-4.1, T=0). Evaluation artifacts. A 155-row consolidated clinician-score CSV from four evaluators (two senior board-certified neurologists, one neurology trainee, and one medical informatician) across five evaluation dimensions. LLM-as-judge results in three configurations: v1 blind, v2 citation-aware, and v3 five-dimension across three providers. A 15-case counterfactual safety experiment with global outcome summary. Citation audits. Type-I PubMed verification of all 198 unique cited PMIDs. Type-II per-claim natural-language-inference audit on 200 stratified (claim, PMID) pairs (1.0% explicit contradiction, 99.0% non-contradiction). Headline numbers reproducible from this bundle 100% Type-I PMID verifiability (198/198 cited PMIDs verified against PubMed E-utilities) 99.0% Type-II non-contradiction (95% CI 97.5–100.0%; n=200 NLI audit) Three-arm clinical feature coverage: HEG-TKG 0.717, Vanilla 0.743, Guideline-RAG 0.688 (Mann-Whitney p > 0.10) Provenance Gap: HEG-TKG 0.366 vs Vanilla 0.743 / Guideline-RAG 0.688 Clinician D1 Verifiability advantage replicated across three independent neurologists (Δ = +1.65, +0.67, +1.30; all BH-significant) Inter-rater reliability between two senior neurologists on D1: Spearman ρ = 0.79, ICC(2,1) = 0.62, quadratic κ = 0.60 Counterfactual safety: 80% parametric resistance (12/15 cases), 100% detectability via citation trace Bundle structure Extract heg-tkg-paper1-data-v1.0.0.zip: scenarios/ - 36 clinical scenarios (Python source-of-truth + JSON dump) kg_snapshots/{dmd_bmd, mg_lems, cidp_gbs}/ - nodes, edges, triplets, Cypher import, statistics clinical_outputs/{pair}/ - 12 scenarios × 3 arms per pair, plus per-pair metrics, citation verification, sealed evaluation key clinician_evaluation/ - anonymized scores CSV plus inter-rater statistics llm_judge_results/ - three judge configurations across three providers counterfactual_experiment/ -15 cases plus global summary citation_audits/type1_pubmed_verification/ - Supplementary S4 citation_audits/type2_nli_audit/ - Supplementary S4.1; full pipeline (extractor, sampler, NLI runner, aggregator), 105 cached PubMed abstracts, per-row GPT-4.1 verdicts; v1 first-pass results retained under archive/ for transparency rarebench_scope_check/ - RareBench scope-mismatch finding (Supplementary S33) Reproducibility All random seeds documented (seed=42 for bootstrap and stratified sampling; bootstrap 10,000 resamples) All model versions and API parameters specified Source code at https://gitlab.sdu.dk/screen4care/heg-tkg/-/tree/paper1-submission-v1.1.0 Anonymization and ethics All clinical scenarios are synthetic vignettes; no patient-derived data Clinician scores de-identified at row level (codes C1, C2, C3, MI); evaluators consent to identification in the manuscript Author Contributions section No protected health information (PHI) anywhere in the bundle License CC BY 4.0 for data and documentation; MIT for accompanying code (citation_audits/type2_nli_audit/*.py). Citation If you use this dataset, please cite both this Zenodo record (DOI 10.5281/zenodo.19763337) and the underlying paper (DOI to be added on publication).

提供机构:
Zenodo
创建时间:
2026-04-25
二维码
社区交流群
二维码
科研交流群
商业服务