遇见数据集

Synthetic electronic health record datasets from 16 generators benchmarked on NHANES 2015 to 2016

收藏
Zenodo2026-05-29 更新2026-06-05 收录
官方服务:

资源简介:

A corpus of synthetic adult electronic health records derived from the National Health and Nutrition Examination Survey (NHANES) 2015–2016 cycle, together with evaluation artifacts produced by the PRISM-EHR pipeline. The source covers 4,809 adults across 21 clinical concepts (demographics, vitals, laboratory panels, 13 diagnosis flags), binarized at clinical thresholds into 38 indicator columns. Sixteen generators were refit on each of 10 frozen 70/30 splits, yielding 160 paired synthetic datasets. Generators span an independence baseline, six classical models (BN, DT-kNN, SCART, MoB, LRC, GCop), three deep latent and adversarial models (VAE, MedGAN, CORGAN), a diffusion model (DDPM), and four GPT-2 small language models (82M, 124M, 355M, 774M parameters). Every synthetic dataset is evaluated against the held-out test half on three fidelity measures (DWD, CWC, LCA) and four disclosure-risk measures (AIR, MIR, MIDR, NNAA), with 95% percentile bootstrap intervals and paired significance tests. This deposit contains the preprocessed cohort, frozen splits, synthetic outputs, and aggregated evaluation results. The pipeline source code is available separately on GitHub and Zenodo (see Related Identifiers). Accompanying: Angeles, S., Liang, M., Burnison, M., Guillen, J., & Vardhan, M. PRISM-EHR: a synthetic electronic health record corpus and evaluation pipeline benchmarked on NHANES 2015 to 2016. Scientific Data (under review).

提供机构:
Zenodo
创建时间:
2026-05-29
二维码
社区交流群
二维码
科研交流群
商业服务