Synthetic electronic health record datasets from 16 generators benchmarked across four clinical cohorts
收藏资源简介:
A corpus of synthetic patient records derived from four clinical cohorts, with evaluation artifacts from the PRISM-EHR pipeline. The cohorts are NHANES 2015-2016 and 2017-2018, the eICU Collaborative Research Database (US multi-centre critical care), and the Zigong Fourth People's Hospital heart-failure cohort (China). Each cohort is reduced to a fixed binary schema and complete-cased. Cohort sizes are 4,809 adults (NHANES 2015-2016), 4,479 adults (NHANES 2017-2018), 60,494 ICU stays (eICU), and 2,008 hospitalisations (Zigong). Sixteen benchmark entries (a Train passthrough plus 15 generators across five categories, Reference, Probabilistic, Neighbour, Deep, Language) were refit on 10 frozen 70/30 splits per cohort, yielding 640 paired synthetic datasets total. Every synthetic dataset is evaluated against its held-out test half on three fidelity measures (DWD, CWC, LCA) and four disclosure-risk measures (AIR, MIR, MIDR, NNAA), with 95% bootstrap intervals, paired significance tests, and a cross-cohort consistency analysis. This deposit contains synthetic outputs and evaluation results only. It excludes raw and row-level real patient data. eICU and Zigong are PhysioNet credentialed-access datasets, and NHANES is public but not re-hosted here. It also excludes pipeline code, released separately (see Related Identifiers). Accompanying: Angeles, S., Liang, M., Burnison, M., Guillen, J., & Vardhan, M. PRISM-EHR. A synthetic electronic health record corpus and evaluation pipeline benchmarked across four clinical cohorts. Nature Digital Medicine (under review).



